Testing content before you commission it: how The Times did it

Testing content before you commission it: how The Times did it

Test a format, headline or commissioning decision against a model of your subscribers in about a minute. How The Times did it, and where it stops working

Readers are not reliable witnesses to their own taste. On Facebook, The Atlantic collects about 27 mentions for every one the National Enquirer gets. In the circulation figures, and in the number of people who actually search for each title, the two are level. Seth Stephens-Davidowitz uses that gap in Everybody Lies to make a plain point: what people say they read is the version they want to be seen reading, and it can sit a long way from what they actually open.

For a commissioning desk, that gap costs money, because most editorial research asks readers to predict their own behaviour, and readers are generous with the highbrow answer. Two problems then stack on top of each other. Readers misreport, and the survey that might correct for it lands weeks after the decision was made. Content testing with a synthetic audience takes aim at both. You put the actual thing, a headline, a format, a trailer cut, in front of a model of your real subscribers, and read how each segment reacts in about a minute, before you commission anything.

What can a publisher test with a synthetic audience?

All the decisions that get waved through in a commissioning meeting because there was never time to research them.


Decision

How it gets made now

What you can ask instead

Whether to commission a series

Editorial judgement plus last quarter's traffic

Which segments would follow this to episode three, and what would make them stop

What to call a podcast or newsletter

A shortlist argued in a room

What each name signals, and to which reader segment

Which format a story should take

Whatever the desk usually does

Whether this audience wants the explainer, the video or the long read

A subscription price or tier change

Benchmarking against competitors

What each segment expects to pay, and what makes it worth that

A paywall or registration change

An A/B test after it ships

Which segment walks away, before it ships

Whether a commercial partner reaches your readers

Two audience decks side by side

Where the two audiences genuinely overlap

An AI-assisted product feature

An assumption about reader appetite

Which readers want it, and which quietly consider it a downgrade

Why is this different for publishers specifically?

Your decision window is shorter than your research window, and no amount of hustle closes it.

The desk decides this week. A reader survey comes back in three to six weeks. By then the story has run and the research arrives as a post-mortem. That is why so much editorial research explains what already happened instead of shaping what happens next. The research is sound. It just arrives after the decision.

Publishers also sit on the opposite problem to most businesses. You are not short of audience data. You have subscription records, engagement data and reader panels, years of it. What you are short of is permission to keep asking. Every survey spends a little of your readers' goodwill, and that well is not bottomless. The Times had 642,000 subscribers and a hard limit on how often it could go back to them.

A synthetic audience takes the questions so the real panel does not have to, which keeps the goodwill for the moments that genuinely need a human on the record. There is more on why that well is running dry in why customers stopped answering surveys.

How did The Times use this?

It ran as a shared research layer, with three teams each pointing it at their own problems, rather than one workflow imposed from above.

Product stopped burning weeks on panel work and started getting reactions in seconds. One test is the reason people started paying attention: loyal readers were wary of AI features, while the readers The Times most wanted to win over were asking for exactly those features, summaries, AI explainers and video. Two opposed answers living inside one audience. Average them and you get a lukewarm number that describes nobody, and it would have set the roadmap. The split set it instead.

Marketing ran CRM optimisation, copy testing and campaign evaluation without rationing questions. The one people quote back: four candidate names for a new business podcast, where the winner, The Business, tested fine with the general reader panel and landed much harder with the exact demographic the show was chasing. The same shortlisting logic sits in message testing with synthetic audiences.

Commercial used it to sense-check partnerships before signing them, asking whether a partner's audience actually overlapped with theirs rather than taking two decks at face value. That evidence had not existed before, and it let them put money behind the deals that reached real readers.

Same team, same budget, roughly ten times the research volume across the three functions. The full story is in The Times case study, and the tool they built on top of it is covered in The Times' synthetic audience tool built on Electric Twin.

"We went from rationing research to running it on demand. The same team, ten times the output, and we could prove it matched our real audience."

Chris Courtney-Smith, Director of Data Operations, The Times

But is it accurate on a real subscriber base?

Fair question, and the honest answer is a number. The Times held back 10% of its known-trusted reader panel, built the synthetic audience from the rest, and scored the audience's answers against the slice it had hidden. The match came in at 92%, next to roughly 93% for the traditional methods they were already trusting.


Measure

Result

The Times synthetic audience, against its own 10% holdout

92%

The traditional methods The Times was already using

Around 93%

Real people agreeing with themselves on a repeated question

Around 94%

Electric Twin on 1-MAE, the industry measure

Up to 96%

Electric Twin on NDAM, the stricter measure

Up to 92%

An ungrounded frontier model on NDAM

65%

That result is what took this out of the insight team and into everyone else's meetings. It is also the number you should refuse to buy without: not a vendor's internal benchmark, but a score on your own subscriber data, with the holdout held back by you. The measures and the protocol are in what a synthetic audience is and how it works.

Can you test the actual thing, not just the idea of it?

Yes, and for a commissioning call that is the whole point. This is the stated-versus-revealed gap from the top of the page, closed from the other side. Rather than asking someone what they think about AI explainers, you show them the actual explainer and read the reaction. Reacting to a cut of a trailer is not the same as holding an opinion about the trailer's subject, and it is the reaction you want.

We validated our audiences against real human reactions to real assets: 18 images and 6 videos of up to three minutes, measured against a benchmark of 4,000 people. None of the opinions those people held about the assets went into building the audiences, so the audiences reacted cold rather than repeating something they had been told. The predictions tracked the real reactions closely, including on which of two assets won a straight head-to-head, which is the shape almost every format decision takes. Image and video work runs through creative testing, and the same platform figures hold, up to 96% on 1-MAE and up to 92% on NDAM, always against real answers held back from the model. The full method is in how accurate synthetic audiences are.

The honest line for a newsroom: this is strong enough to shortlist a thumbnail, a trailer cut or a format treatment down to the ones worth backing. It is not the final word. Narrow with it, then commit with editorial judgement.

What did adoption actually take?

More than a year in, the thing The Times kept flagging was cultural, not technical.

Roll this into a research team and it splits people in two. Some see their capability multiply overnight. Others see a threat to the thing they are trusted for. That divide does not close because you waved a good accuracy score at it. It closes with visible human oversight, repeated validation against real responses, and a message that never wavers: this extends what researchers do.

The second lesson was about proof. Trying to show that any single editorial decision came out better is close to unprovable, so instead they counted what they could count, research volume and speed to delivery, and put those in front of leadership. That was less romantic than proving decision quality, and far more persuasive.

What does a publisher need to start?

Real data about the readers you want to model: subscription records, engagement data, old survey waves, reader panel responses. We build the audience from that and validate it against a held-out slice of the same data before you rely on a word of it.

The sector walkthrough, from subscription pricing to the greenlight meeting, is on Electric Twin for media and entertainment. Pre-launch format work lives on concept and idea testing.

What should you test first?

Start where being wrong is cheap to discover and expensive to ship. Three phases, in order, because you have to earn the right to test the decisions that carry weight.

First, test something you already know the answer to. Re-run a question your last reader survey answered, and compare. This feels like a waste and is not. It is how the insight team walks into the newsroom with a number it can defend when someone asks why anyone should believe a synthetic audience. That first piece of internal proof does most of the work.

Second, test the decisions nobody bothers to research. Names, subject lines, format calls, newsletter positioning. Low stakes, high frequency, currently decided by whoever talks loudest. Any evidence is an upgrade, and nobody feels their job is under threat. This is also where the tenfold jump in volume comes from: questions that were never in the research plan.

Third, bring it to the decisions with money on them. Subscription tiers, paywall changes, a series commission, a partnership. By now it has a track record inside the building, which is what makes the answer usable in a room where budget is being signed off.

Get the order wrong and you start the credibility fight before you have any evidence to win it, which is how good pilots quietly die.

One more thing, on ownership. The Times ran it as a shared layer across product, marketing and commercial, each team pointing it at their own problems. The moment a single team owns it, it becomes a service-request queue, and you have rebuilt the bottleneck you were trying to escape.

Where does this fall over?

A thin or brand-new audience. Launching into readers you hold no data on gives the model nothing to stand on. Do the primary research first.

Anything that needs a name on the record. Regulated claims and legally mandated testing need real, identifiable people. No exceptions.

A replacement for the reader panel. The panel is now your calibration layer, and skipping the periodic real-world checks lets the audience drift away from your actual readers without ever announcing it. This is a documented failure mode. Google Flu Trends predicted flu from search data, impressively at first, then drifted so far that it became a cautionary tale and was retired. A synthetic audience that is never re-checked goes the same way, quietly.

A stand-in for editorial judgement. It tells you what a modelled audience would say. Whether that should change what you commission is your call, and it should stay your call. A title that commissions purely to a predicted preference has stopped having a point of view, and readers clock that faster than they clock any format change.

Frequently asked questions

Can a synthetic audience be built from subscription data? Yes. Subscription records, engagement data and existing survey datasets are all usable, and the audience is validated against a held-out slice of that same data.

Can it react to an actual video or image? Yes, and that has been validated against a 4,000-participant human benchmark across 18 images and 6 videos of up to three minutes.

How is this different from analysing the audience data we already have? Analytics tell you what readers have already done. A synthetic audience answers the question you have not asked yet, including about formats and products that do not exist.

Does this work for commercial and advertising teams? Yes. The Times used it to model audience overlap before committing to partnerships, and put spend behind the deals that actually reached their readers.

Can we test commercially sensitive ideas? That is one of the best reasons to use it. A pre-launch format or a competitor comparison stays inside the building instead of going out to a panel.

How often does it need re-validating? Periodically, against fresh real responses. Readers move, and re-validation keeps the model moving with them.

Does this kill off the reader panel? No. You draw on it less often, which tends to make it last longer. The panel stays essential as the calibration layer.

Stop guessing.

Start predicting

Stop guessing.

Start predicting

Stop guessing.

Start predicting