AI for Media: Editorial Automation with RAG and Human Oversight
We automate what repeats. The AI looks for what, on that particular day, is different.
BigLearn proprietary technology, at mature proof-of-concept stage. This is not the account of a deployment inside a media group: it is the platform we built, tested end to end, ready to be implemented and tailored to a newsroom in an estimated 3 to 9 months depending on the number of pipelines, sources and integrations. We publish no audience or production metrics — they would be numbers we do not yet have. What is here is the architecture, the editorial criteria and the boundaries: what the platform automates, what needs approval, and where it deliberately does not go.
What is artificial intelligence for media?
It is automating the monitoring, collection, validation, preparation, publication and distribution of recurring news, with journalists and editors keeping control. The solution combines previously approved sources, scraping and APIs, structured data, RAG, language models, agents, CMS integration, SEO, AEO, multichannel distribution and an executive dashboard.
The aim is not to replace journalism. It is to free newsrooms from mechanical work, publish faster the information the public is actively looking for, and create time for investigation, interviews, analysis and original sources.
It works when three conditions hold at once: we know what to look for, when to look, and where to look. That covers lottery draws, football results, Euribor, inflation, unemployment, house prices, fuel, weather, tourism, stock indices, international markets, Brent, gold and silver.
The flow, from event to measurement
event and source calendar → monitoring → scraping or API → validation → structured data → historical comparison → RAG → relevance analysis → demand analysis → LLM → human oversight → CMS → update → distribution → measurement → editorial decision
The language model is not used as a factual source. The source provides the fact, software validates it, and the AI helps identify and communicate what it means. It is the same boundary we apply in cybersecurity by design: the model interprets and proposes, deterministic components verify.
The problem is not the story, it is the repetition
How long does it take to write up a lottery draw? Or explain that Euribor fell? Or update the fuel price forecast? Individually, minutes. The problem is doing it several times a day, every day — morning, market close, after the matches, at night, at weekends, on public holidays. Next day it starts again.
These are stories of real public interest and often little interest to write. For the reader, Euribor or fuel prices hit the mortgage and the household budget directly. For someone who has written the same type of story hundreds of times, collecting the same fields again is mechanical work. That gap is exactly what makes them good candidates for automation.
The platform knows when to look, and where
Many editorial events have a known cadence. For each topic the platform stores event + source + URL or API + expected time + frequency + fields to collect + validation rules + history + relevance criteria + editorial rules.
The newsroom stops depending on someone remembering that the inflation figures land that morning. And knowing the source matters as much as knowing the hour: the application does not ask a model what inflation is — it fetches the new figure from the source configured for that pipeline.
Speed is audience too
Many of these stories happen outside newsroom hours. A draw at night. A match ending at 11pm. Wall Street closing after the working day. Asian markets running while Europe sleeps.
The reader wants the answer immediately, and there is a concentrated spike of demand right after the event. Arriving ten or twenty minutes earlier makes a measurable difference. Speed stops being an operational detail and becomes an editorial and audience variable.
If it were only filling fields, a template would do
A template can write that a figure went up. AI adds value when it compares the new data against context: the previous value, the forecast, the same period last year, highs and lows, runs of rises or falls, position in a table, unusual patterns, and the group's own recent coverage.
The event, the source and the structure are predictable. What the new data means is not.
Publish by relevance, not by calendar
New data does not oblige a new story. If fuel rises by one cent, there may be no story. But if small rises repeat for three months, the accumulation becomes one — and the angle shifts from «fuel up one cent» to «fuel has risen for weeks and is now significantly higher».
The platform can start by producing on every scheduled date and evolve into editorial monitoring: preparing or publishing only when the newsroom's criteria are met. Monitoring everything does not mean publishing everything.
Where this applies
Inflation and economic indicators
If inflation moves from 2.9% to 3.3% against a 3.0% expectation, the best headline may not be «inflation at 3.3%» but «inflation accelerates faster than expected». Some weeks bring inflation, unemployment, GDP and trade almost back to back.
Euribor and mortgages
Daily interest in the three, six and twelve month rates. A general title explains the impact on the monthly payment; a business title gives basis points, monthly average and historical comparison. Same figure, different treatment.
Housing and regional angles
One source with enough granularity produces national, regional and local stories. Without automation, turning a dataset into ten pieces means repeating the work ten times.
Football and factual sport
The result is factual but rarely the whole story. A leader dropping points against the bottom club can matter more than a predictable rout. After the whistle the reader's question is: what changed?
Lottery draws and sensitive data
A wrong number makes someone believe they won millions. Source, date, draw identifier, count of numbers, permitted ranges, timestamp and difference against the previous draw are all validated before anything is written.
Markets, Brent, gold and silver
Markets closing daily does not mean a story daily. But if Lisbon falls alongside Frankfurt, Paris and Wall Street there is a global frame; if it rises while the others fall, the divergence is the story.
RAG: the archive becomes a knowledge base
Before generating, the platform retrieves related material published by the group itself — last week's fuel piece, the match preview, the analysis by the group's own economics correspondent. The publication starts reusing its own journalism intelligently.
At 11pm it would be inefficient to ask a journalist to open several past pieces, compare years of data, build internal links and prepare versions for different brands. Paradoxically, a story written by hand under deadline pressure can come out more mechanical than an assisted draft — because the draft arrives with history and context already gathered.
Human-in-the-loop: the newsroom sets the autonomy
Early on: draft → journalist → approval → publication. On a stable, proven pipeline: validation → publication → notification. A story with sensitive data requires prior approval; a simple update from a stable source can publish and be reviewed afterwards.
The journalist does not have to enter a new dashboard to approve each piece — it arrives on the channel the team already uses. The newsroom always decides what is automated, what needs approval, which sources are accepted, and when the automation should stop.
SEO and AEO: answer first, contextualise after
Publishing early puts the piece inside the interest peak, but that is not enough: it has to use the language the public uses. An editorial headline can coexist with an explicit variant that makes the entity, the result and the consequence immediately legible to both search engines and readers.
For AEO, the order that works is: direct answer in the headline or lead; the main figure and its source; impact on the reader; comparison with the previous value; historical context retrieved by RAG; frequently asked questions; internal links. Someone asking how much fuel will rise should find the number in the first sentence.
One live story, several lives, the same URL
Speed and depth do not have to compete. A draw story can evolve from the numbers, to the prize breakdown, to news of a national winner. Keeping the same URL avoids fragmenting audience, links and relevance across near-duplicate pages. The story stops being an article and becomes a live editorial asset.
This is not a content farm
An «intelligent editorial factory» does not mean producing thousands of articles because generating is cheap. It means monitoring a lot and publishing what deserves it. A good candidate combines public interest, an identified source, a known cadence, new data, the ability to validate, context and editorial relevance. When relevance disappears, the platform keeps watching without publishing.
And it does not replace investigation, interviews, reporting, original sources, sensitive political analysis or deep editorial judgement. The first story can be automatic. The big story stays human.
Technology
Generative AI · LLMs · Retrieval-Augmented Generation · AWS Lambda · Python · Web scraping · APIs · Parsing · Structured data · Historical data · Google Trends · Search Console · Search intent analysis · SEO · AEO · Google Discover · AI agents · Prompt and context engineering · Human-in-the-loop · CMS integration · Social media automation · Workflow automation · Executive dashboard
Frequently asked questions
What is artificial intelligence for media?
It is applying AI, automation, data and systems integration to editorial work. It can monitor sources, validate information, prepare stories, retrieve archive context, optimise content for SEO and AEO, distribute across channels and measure results.
Can AI produce news automatically?
Yes, particularly where the data is structured, the source is known and validation rules are clear. The newsroom decides whether to require human approval before publication or to use post-publication review on pipelines that have proved reliable.
Which stories can be automated?
Those with a known cadence and an identified source: official results, factual sport, lottery draws, Euribor, inflation, unemployment, housing, fuel prices, weather, stock indices, international markets, Brent, gold, silver and tourism.
How is data accuracy guaranteed?
Facts come from previously approved sources and pass deterministic validation — date, format, permitted ranges, field count, timestamp, event identifier and comparison against history. The language model does not replace the factual source.
What is RAG in a newsroom?
It is retrieving relevant material from the publication's own archive before generating new text. The story can then carry history, context, internal links, earlier analysis and journalism produced by the group itself.
How does editorial automation improve SEO?
It helps publish during the demand peak, maintain regular coverage, build internal links, update the same URL instead of fragmenting audience, reinforce topical authority and match headlines to real search intent.
How does editorial automation improve AEO?
It structures content to answer the reader's question first and organise context into blocks answer engines can interpret. Someone asking how much fuel will rise should find the number in the first sentence, not the fourth paragraph.
Does AI replace journalists?
No. It automates repetitive tasks and prepares drafts, while journalists and editors keep control of sources, criteria, tone, approval, interpretation and publication. The time freed goes into investigation, interviews and reporting.
Can the solution work across several brands and regions?
Yes. The same raw material can generate national, business, regional, specialist and social versions, each following its own brand's editorial rules. Collection happens once; what changes is the reuse.
Can it integrate with the CMS and newsroom tools?
Yes. The architecture integrates with a CMS, internal channels such as Teams, WhatsApp or email, APIs, dashboards, analytics platforms and social channels, according to the systems the group already runs.
Is this not a content farm?
No. A content farm publishes a lot because generating is cheap. Here you monitor a lot and publish what deserves it — when relevance disappears, the platform keeps watching without publishing.
Who did this work
BigLearn is a Portuguese artificial intelligence consultancy, founded in 2017 and based in Lisbon.
We do not start by asking «how can we use a language model?». We start by working out what information exists, when it appears, where it appears and who does the manual work today. This case comes out of our AI agents and business automation and SEO and digital transformation work.
The other case studies are published under the same rule: client anonymised, verifiable figures, and the nature of the document stated up front.
Other case studies
Tourism, Hospitality and HORECA
- Event Configurator for a Hotel Group
- AI for Hotels: a Guest Journey Agent
- AI and Marketing to Win More MICE Leads for Hotels and Event Venues
- AI for HORECA: Hours, Allergens and Orders
- Digital Gift Vouchers for a Hotel Group