Synthetic respondents, personas, panels: what's the difference?

Synthetic respondents, personas, panels: what's the difference?

Synthetic respondents is the industry's term for modelled participants built from real data; we build them into a synthetic audience. Here is the difference that matters, the failure modes to watch for, and a validated result

Two vendors can both say "synthetic respondents" and mean opposite things. One has modelled a real population from survey data and scored its answers against responses it never saw. The other has typed a paragraph describing a customer and asked a chatbot to play along. Same words, no shared standard underneath them. Ipsos, Qualtrics and NielsenIQ are all selling versions now, so the vocabulary is only going to get more crowded and less precise.

Two questions cut through all of it. What is the thing built from? And has anyone scored its answers against real answers it was never shown? Everything else is packaging.

What are synthetic respondents?

A synthetic respondent is a modelled participant, built from real survey or behavioural data, that answers a research question the way a member of a defined population would. One on its own is not much use. Model the whole population and you get a distribution instead of a single manufactured opinion, which is why we work at the level of a synthetic audience rather than a lone respondent.

Strip away the vocabulary and two questions separate a usable method from a useless one:

What is it built from? Real response data about a specific population, or a sentence somebody typed into a prompt.

Has the output been scored? Against real answers held back from the model, or not at all.

Everything else, including the format it shows up in, sits downstream of those two. And since the formats get sold as though they were methods, they are worth pulling apart.

Respondents, personas, panels, AI focus groups: what's actually different?

Vendors use these loosely enough that you can end up comparing a method with a format and never notice. Here is the taxonomy that matters.


Term

What it actually is

Built from

Can it produce a distribution?

Synthetic persona

A described character, often handed to a general model to act out

A description someone wrote

No. One voice.

Synthetic respondent

One modelled participant answering as a member of a defined population

Real response data about that population

Only as part of a full audience

Synthetic audience

The whole modelled population, questioned together and read by segment, which is what we build

Real survey or first-party data, held as a reusable audience

Yes, and comparably across waves

Synthetic survey

A format. A structured instrument run against synthetic respondents

Whatever the respondents were built from

Depends entirely on the respondents

AI focus group

A format. A moderated conversation with synthetic respondents

Whatever the respondents were built from

Not its purpose; it probes reasons

Synthetic data generation

A different discipline. Statistically similar records, usually for privacy or model training

A real dataset's statistical structure

Not a research method for asking questions

Three things fall out of that table, and they are the ones buyers keep missing.

The format tells you nothing about the quality. A synthetic survey and an AI focus group run against ungrounded personas are both junk. The same two formats run against respondents built from real data and scored against a holdout are both usable. People evaluate the interface and forget to ask what is underneath it.

Persistence is worth more than it sounds. A panel you can go back to lets you compare this month with last month, because the respondents are the same respondents. Regenerate the audience from scratch each time and wave-on-wave comparison quietly becomes meaningless. A surprising number of tools work exactly that way.

Synthetic data generation is a different job. It turns up in the same search results and solves a privacy and volume problem, not a research question. Confusing the two is the most common mix-up in this category.

How are synthetic respondents built?

Four stages: grounding, construction, questioning, validation. The last one decides whether the first three counted for anything.

Grounding. Real survey data from leading global survey companies, your own datasets, or both. Our published evaluations run across internal datasets spanning UK and US populations. The audiences are built using large language models that take on characteristics drawn from that survey data, so the difference from a general-purpose model is the grounding and the construction, not the absence of a model.

Construction. We build the audience from what the data shows about beliefs, preferences and behaviour, one individual at a time. Building at that level is what makes a distribution possible in the first place. Model the population average instead and you get a single voice, and a single voice cannot disagree with itself.

Questioning. Surveys, focus groups, message tests, image tests. Answers come back in seconds.

Validation. We hold a portion of the real data back and never show it. The respondents answer questions drawn from it, and we score those answers against what the real people actually said. The evaluation questions never touch the build, so a good score cannot come from memorisation.

The full mechanics are in what a synthetic audience is and how it works.

Fine, but is it accurate?

Most of the category answers this in the abstract. Here are numbers instead.

We report two measures. Up to 96% on 1-MAE (one minus mean absolute error), the industry benchmark, which measures how closely the modelled answer matches the average real answer. Up to 92% on NDAM (normalised distribution accuracy measure), which scores how closely the model predicts the shape of the entire distribution.

Why both? Because 1-MAE gets more generous as a question gains answer options, so the same method flatters itself on a ten-option question next to a two-option one. NDAM holds its meaning whatever the question looks like, which is why it is the number we ask to be held to.



1-MAE

NDAM

Full name

One minus mean absolute error

Normalised distribution accuracy measure

What it scores

How close the average answer is

How close the shape of the whole distribution is

As answer options increase

Gets more generous

Holds its meaning

Tells you about outliers

Nothing

Who dissents, and where opinion splits

Electric Twin

Up to 96%

Up to 92%

An ungrounded frontier model (GPT-5)

Not comparable

65%

The volume behind those figures: 50,000 or more evaluations against real survey responses, across 155 countries, benchmarked against eight traditional methods. Both come from held-out real answers rather than a vendor demo, and both get re-measured on your own data at onboarding. Accuracy has also climbed by an average of 20% over the period we have tracked it, so treat any figure as a floor. The number worth demanding from any vendor is the result on your data, with a holdout you controlled, not the one on the front of the deck.

Show Image Distribution accuracy on NDAM, against a frontier model with no grounding and against the rate at which real people agree with themselves. Source: Electric Twin accuracy methodology.

Do the answers actually sound human?

There is a published test, and the result cuts both ways, which is the interesting part.

We put blind comparisons of a synthetic discussion group against a real one, across UK issues, AI in banking and football club ownership, in front of three AI judges. In 6 of the 9 judgments, the synthetic group was rated the more authentic of the two.

You can read that as a win, and it makes a good headline. The more useful reading is a warning. Models are trained to be articulate. Real people ramble, contradict themselves, and say things that make them look bad. A synthetic transcript that reads better than a real one may simply be reading too well, which is the polished-rationalisation failure mode showing up in a formal test. So a fair question for any vendor, us included: has your qualitative output ever been mistaken for real, and do you treat that as a pass or a red flag?

What's the realistic ceiling?

About 94%, and it is the single most useful number to walk into a procurement meeting holding.

Real people disagree with themselves. Ask someone the same question twice in one survey and they will change their answer roughly six times in a hundred. So no method has a clean 100% to be measured against, human panels included. A vendor quoting 99% should worry you: at that level they are either measuring something forgiving or claiming a consistency real humans do not have.

There is a related wrinkle: where you ask changes the answer you get. We published research on this in March 2026, Where You Ask Matters: Platform Effects of Online Surveys and AI-Based Human Subject Simulation, and the plain-English version is in the same survey run across seven panels.

What are the real failure modes?

Each has a specific cause, and knowing the cause turns "sounds risky" into a question you can put to a vendor.

Stereotype completion. Ask a model built on demographic labels what a 34-year-old woman in Leeds thinks and it hands back the most statistically available answer for that description, a stereotype delivered with total fluency. Walter Lippmann named this trap in 1922, when he coined the modern sense of the word "stereotype": most of the time, he wrote, we do not first see and then define, we define first and then see. A model keyed to labels does the same, in reverse order from a real person. Cause: building respondents from which boxes someone ticks rather than what they actually said. Test: ask about a segment you know cold and see whether the answer contains anything you did not already assume.

False consensus and flattened variation. A model can nail the average and still crush the spread, reporting 80% agreement where the real figure was 55%. This is worse than it sounds, and there is a theorem that says so. In The Model Thinker, Scott Page states the diversity prediction theorem as an identity: the error of a group of predictions equals the average individual error minus the diversity of those predictions. Diversity is not decoration around the average. It is a measurable part of the accuracy. So a model that flattens the spread is not merely less colourful, it is less accurate, and a distribution-level measure exists to catch exactly that. A method can look strong on 1-MAE while getting the shape of opinion badly wrong.

Social desirability and polished rationalisation. Models are trained to be helpful and articulate, so they produce reasoning tidier than real human reasoning. Real respondents contradict themselves and give unflattering answers. Test: go looking for answers that make your product look bad. If there are none, be suspicious rather than relieved.

Overclaiming from apparent precision. A synthetic result arrives formatted, complete, and free of the missing data that would flag uncertainty in a real survey, so it looks more precise than it is. The discipline is simple: state the validation next to the number, every time.

Hallucination from weak grounding. Thin input produces confident output anyway. The model does not degrade gracefully, and it will not tell you the data was thin.

Four of those five are the same failure wearing different outfits: an average passing itself off as an answer. Respondent-level construction and a distribution-level measure are what separate a method from a demo.

How does this compare with what the incumbents offer?

The established research houses have all moved into this space. Rather than characterise anyone else's methodology, the sharper move is to know which questions separate a real offer from a repackaged one. Take these to any vendor, us included.


Ask the vendor

What a good answer sounds like

What to worry about

What is the audience built from?

Named real datasets, and our own data if we provide it

"Advanced AI models" with no dataset named

Which accuracy measure is that, and what does it score?

A named measure with its behaviour explained

A bare percentage with no measure attached

Will you validate on our data before we buy?

Yes, on a holdout you control

An internal benchmark only

Can I see the full distribution as well as the average?

Yes, with the dissenting segments identified

One summarised voice per question

What will you refuse to do?

A specific list of question types

"It works for everything"

How is drift monitored?

Periodic re-validation against fresh real responses

No answer, or "the model keeps learning"

That fifth row is the most diagnostic question on the list. A vendor with no refusals has not tested the edges of its own method, so you will end up finding them instead, usually mid-project.

What does a validated result look like in practice?

A real one, rather than a worked hypothetical.

The Times held back 10% of its known-trusted reader panel, built a synthetic audience from the rest of a 642,000-subscriber base, and scored the audience's answers against the slice it had kept back. The result was a 92% match, against roughly 93% for the traditional methods already in use.

That single result is what opened the doors. Synthetic research went on to run across three functions, product, marketing and commercial, delivering ten times the research volume on the same team and no extra budget. The distribution earned its keep as much as the headline did: long-standing subscribers and prospective ones wanted opposite things from the same product question, and an average would have reported a preference neither group actually held.

"We went from rationing research to running it on demand. The same team, ten times the output, and we could prove it matched our real audience."

Chris Courtney-Smith, Director of Data Operations, The Times

The full account is in The Times case study.

What did adoption actually require?

Less than the accuracy debate suggests, and more of one thing teams underestimate: it is a cultural change, not a technical one. Dropped into a research function, it splits people into those who see their capability grow and those who see a threat to what they are trusted for. That gap closes with visible human oversight, ongoing validation against real responses, and a consistent message that this extends what researchers do. The Times also measured adoption by research volume and speed to delivery rather than by trying to prove any single decision improved, which is close to unprovable. Expect the credibility conversation before the procurement one.

When should you still use human respondents?

We publish these rather than bury them.

Anything on the record. Regulated claims, clinical studies and legally mandated consumer testing need real people who can be identified and counted.

Anything physical. Taste, packaging in the hand, usability with a device, in-person ethnography.

Anything genuinely novel. A category with no prior behaviour in the data gives a model nothing to build from.

Calibration, always. Traditional research is the yardstick these scores are measured against, so periodic real-world validation is what keeps synthetic respondents accurate. Take it away and the method stops being trustworthy, slowly and invisibly.

And where the audience data itself is thin or newly formed, primary research comes first.

How do synthetic respondents fit into a research programme?

Most teams end up using them in all three of these spots.

Before human research, to narrow the field. Nine propositions is too many to field. Run the nine against synthetic respondents, take the surviving two to real ones.

Between waves, to keep a tracker alive. Where the budget only stretches to quarterly fieldwork, synthetic respondents carry the months in between, with each real wave recalibrating the model.

On the questions that never made the programme at all. The niche segment, the confidential comparison, the decision too small to justify a study. This is where most of the volume jump comes from, and it adds to the programme rather than replacing any of it.

Both measures and the published research are on our accuracy methodology.

Frequently asked questions

What's the difference between synthetic respondents and synthetic personas? A persona is usually a description someone wrote, sometimes handed to a general-purpose model to act out. Synthetic respondents are built from real response data about a defined population and scored against answers held back from the model.

Are synthetic respondents representative of real populations? To a measurable degree, which is the only honest form of that answer. We report up to 92% on NDAM against a human self-agreement ceiling of roughly 94%, measured on held-out real responses.

Can synthetic respondents replace a panel? Not entirely, and the reason is structural. Real respondents calibrate the model, so a synthetic programme keeps a real one behind it. What changes is that the panel stops being the only way to get an answer, which also stops it being worn out.

Do synthetic respondents remove sampling error? No. They move where the uncertainty sits, from sampling variance to model fidelity, and fidelity is task-specific. That is why validation runs per audience and per question type rather than as one site-wide number.

How do you know the model isn't just repeating its training data? Because the evaluation data is held out and, in our testing, is private data not available on the internet. A model reproducing training data would fall over on questions it has never seen.

Can we test this on our own data before committing? Yes. Every onboarding takes your real data, builds the audience and runs the holdout evaluation, so the figure you see is measured on your audience rather than ours.

Stop guessing.

Start predicting

Stop guessing.

Start predicting

Stop guessing.

Start predicting