Synthetic vs traditional research: A side-by-side, with the numbers

Synthetic vs traditional research: A side-by-side, with the numbers

Traditional research takes weeks and stays the benchmark; a synthetic audience answers in seconds at up to 92% on NDAM. Here is the side-by-side, and which questions belong to each

Ask The Times's readers about adding AI features to the paper and the audience does not give you one answer. Its loyal subscribers were wary of them. The readers it hoped to attract asked for summaries, explainers and video. Average those two positions and you get a middling number that describes no one, and a roadmap pointed at a reader who does not exist.

Traditional research struggles to catch a split like that without a large, expensive sample in every segment. A synthetic audience catches it in about a minute, because it models each real reader instead of reporting a single mean. That is one job among several where synthetic and traditional research divide the work between them. Traditional research recruits real people, takes weeks, and stays the benchmark every synthetic method is scored against. Synthetic research models your customers from data you already hold and answers fast enough that you can change the question and ask again. Most teams we work with run both. What follows is which questions go to each, and the numbers behind the choice.

What is the difference between synthetic and traditional market research?

Traditional research asks real people and reports what they said. Synthetic research models a specific real audience from data those people already gave you, and predicts what they would say. The prediction is then scored against real answers the model was never shown, which is the only reason this comparison can be quantified at all.

The two answer differently shaped questions. Traditional research produces evidence on the record. Synthetic research produces answers fast enough to turn research from a procurement exercise into something closer to a conversation. Neither substitutes for the other, so the useful version of this comparison is about which questions go where.

How do they compare, side by side?



Traditional market research

Synthetic market research

Time to an answer

Weeks, including recruitment and fieldwork

In seconds

Cost per additional question

Fieldwork cost, every time

No additional fieldwork cost

Iteration

Locked once the questionnaire is in field

Change the question and run again

Panel fatigue

Real, and it caps how often you can ask

None, the real panel is left alone

Hard-to-reach audiences

Expensive, sometimes not feasible

Reachable where enough real data exists

Small or niche segments

Sample sizes get too thin to read

Segment-level answers from the same audience

Accuracy on NDAM

Around 94%, the human ceiling

Up to 92%

Regulated and clinical claims

The only option, respondents on record

Not suitable

Physical and sensory testing

The only option, taste, packaging, in-person use

Not suitable

Confidential or pre-launch questions

Exposes the idea to a panel

Stays inside the business

Where is traditional research still better?

None of these is an edge case.

Anything that has to be on the record. Regulated claims, clinical studies and legally mandated consumer testing need real people who can be identified and counted. A vendor telling you otherwise is a vendor to walk away from.

Anything physical. Taste tests, packaging in the hand, usability with a device, in-person ethnography. The question is about a body in a room, and no model reaches it.

Anything genuinely new to the world. A category nobody has encountered leaves no prior behaviour in the data for a model to build from. Thin or newly formed audiences fall in here too: accuracy depends on the depth of real data behind the model, so a shallow audience needs primary research first.

Calibration. Traditional research is the yardstick the accuracy scores are measured against, so periodic real-world validation is how a synthetic audience stays accurate. This is the point most comparisons skip, and it changes how you read the whole choice. The survey programme does not disappear when you adopt synthetic research. It becomes the calibration layer that keeps the model honest.

What can synthetic research do that traditional can't?

Mostly, it answers the questions that never got asked because asking cost more than the answer was worth.

Questions too small for a study. A subject line, a podcast name, a pricing threshold on one tier. These get settled by opinion in a meeting, because commissioning fieldwork for them is impossible to defend. Testing a shortlist like this is covered in message testing with synthetic audiences.

Audiences too niche to reach. A B2B segment of two hundred. High-net-worth individuals. Legislators. Policyholders who have claimed exactly once.

Questions you can't let out of the building. A pre-launch proposition or competitor comparison exposes itself the moment it goes to a panel.

Questions worth asking repeatedly. Once fieldwork cost drops out, a tracker can run monthly instead of quarterly.

The spread, not just the average. This is the one that changes decisions. Darrell Huff made the point in 1954 in How to Lie with Statistics: the word "average" hides which average you mean. He described a neighbourhood where the mean income was $15,000 and the median was $3,500, the same families, two honest "averages", opposite impressions. A single figure can sit between two groups who want opposite things and represent neither. That is the Times split from the top of this page: loyal readers and prospective readers pulling in opposite directions, with the mean pointing at a reader who is not in the room. Averaged research is structurally unable to show you that. Segment-level synthetic answers show it by default. Broad audience labels fail for the same reason, which is the subject of why Gen Z generalisations are bad for business.

What does the accuracy comparison actually look like?

Start with why a comparison is even possible. Most research opinions are never checked against what actually happened, so nobody finds out whether they were any good. In Superforecasting, Philip Tetlock reports the finding that made his name: across a landmark study of expert predictions, the average forecaster was only slightly better than chance, and the forecasters who got better were the ones whose predictions were scored against outcomes, over and over. Keeping score is what turns a guess into a forecast.

A synthetic audience is built to be scored. We split a real dataset, build the audience from one part, hold the rest back, and measure the synthetic answers against real ones the model never saw. That scoring is the only reason this comparison has numbers.

We report two measures. Up to 96% on 1-MAE, the measure the research industry uses. Up to 92% on NDAM, a stricter measure that scores the shape of the whole distribution rather than how close the average landed. Two numbers rather than one, because 1-MAE gets more generous as a question gains answer options.

The figure that puts the rest in context is the ceiling. Ask a real person the same question twice in one survey and they agree with themselves roughly 94% of the time. There is no 100% for any method, human panels included. A synthetic audience at up to 92% on NDAM sits about two points below a ceiling that real respondents also fail to reach. For contrast, a frontier model with no grounding in your audience reaches 65% on the same measure, which is most of what grounding in real customer data buys you.

Show Image Distribution accuracy on NDAM. Synthetic research against an ungrounded frontier model, and against the rate at which real people agree with themselves. Source: Electric Twin accuracy methodology.

We have run 50,000 or more evaluations against real survey responses across 155 countries, and have been benchmarked against eight traditional methods. Accuracy has improved by an average of 20% over the period it has been tracked, so treat the published figures as a floor rather than a fixed property. The full detail is on our accuracy methodology.

What does each one cost?

Traditional research charges per study, and each additional question carries fieldwork cost. Synthetic research is priced on the audience rather than the question, so the marginal cost of the eleventh question matches the first.

That changes what gets asked more than what gets spent. At The Times the saving came back as ten times the research volume across three teams, rather than a smaller invoice. That is usually the better outcome, and always the easier one to defend upstairs.

"We went from rationing research to running it on demand. The same team, ten times the output, and we could prove it matched our real audience."

Chris Courtney-Smith, Director of Data Operations, The Times

What goes wrong when teams adopt synthetic research?

The hard part is cultural, and it was the most useful thing The Times reported after more than a year of running it.

Synthetic research splits an organisation into two groups: people who see expanded capability, and people who see a threat to their credibility. Managing that divide takes more than an accuracy score. It takes ongoing validation against real responses, visible human oversight, and a steady message that synthesis extends what researchers do.

The second lesson was about how value gets framed. Rather than trying to prove that individual decisions improved, which is close to unprovable, The Times measured research volume and speed to delivery instead. That framing was honest, concrete, and something senior leadership could act on. Plan for the credibility conversation before the procurement one. The detail is in The Times case study.

How do you validate synthetic against traditional research?

By holding real data back.

Split a real survey dataset. Build the synthetic audience from one part. Hide the other and never show it to the model. Then score the synthetic answers against what the real people said. Because the evaluation questions never appear in audience creation, a good score cannot come from memorisation.

The Times ran exactly this on a 10% holdout of a 642,000-subscriber base and got a 92% match, against roughly 93% for the traditional methods already in use. That is the comparison worth asking any vendor for: not their internal benchmark, but a result on your data, held back by you.

One caution before you treat any single survey as ground truth. Run the same survey across different panels and the panel itself shifts the answer. We published research on this in March 2026, Where You Ask Matters: Platform Effects of Online Surveys and AI-Based Human Subject Simulation, and the plain-English version is in the same survey run across seven panels.

Which method for which question?

The choice is rarely about the method in the abstract. It is about the question in front of you, and the cost of getting it wrong.


The question

Method

Why

Which of nine propositions is worth developing?

Synthetic

Nine is too many to field. The job is narrowing, not proving.

Does the surviving proposition justify the campaign budget?

Traditional

High stakes, on the record, worth the fieldwork.

What would this niche B2B segment pay?

Synthetic

The real sample is too small to read at any price.

Can we make this claim in market?

Traditional

Regulated. Needs identifiable respondents.

How does category sentiment move month to month?

Synthetic, calibrated periodically

Monthly fieldwork is unaffordable; quarterly real waves keep it honest.

Does the new pack feel premium in the hand?

Traditional

Physical. No model reaches it.

Which segment will object to this change, and why?

Synthetic

Needs the distribution, not the average.

Is our competitor comparison landing?

Synthetic

Confidential. A panel exposes it.

Synthetic research narrows and explores. Traditional research decides and certifies. The teams that get value from both run them in that order rather than choosing between them.

Frequently asked questions

Are synthetic audiences as accurate as traditional research? Close, and measured against it. We score up to 92% on NDAM, where the human ceiling, the rate at which real people agree with themselves on a repeated question, sits around 94%.

Can synthetic research be evidence for a board decision? Where the validation is on the record, yes. A holdout score against your own customer data is checkable in a way a research opinion is not. Regulated claims still need human respondents.

Do synthetic audiences replace focus groups? Not as a category. They replace specific uses of one. Early exploratory rounds, where a team is narrowing options and does not yet know which question matters, are the rounds a synthetic audience does faster. Once the question is identified and the stakes justify evidence on record, real respondents earn their cost.

Do you still need a survey programme? Yes. Periodic real-world validation keeps a synthetic audience calibrated, so the survey programme becomes the calibration layer rather than the only source of answers.

Which is cheaper? Synthetic research per question, by a wide margin, because there is no fieldwork. The saving usually appears as more research rather than a smaller budget.

Does panel fatigue affect synthetic audiences? No, and that is a large part of why organisations with a valuable first-party panel adopt it. Questions go to the synthetic audience and the real panel is preserved. More on falling participation in why customers stopped answering surveys.

Is this the same as an AI focus group? An AI focus group is one format a synthetic audience supports. What matters is not the format but whether the underlying respondents are grounded in real data and scored against held-out answers. The mechanics are in what a synthetic audience is and how it works

Stop guessing.

Start predicting

Stop guessing.

Start predicting

Stop guessing.

Start predicting