Accuracy
As accurate as traditional research, with published proof so that you can check the claims yourself
Up to 96%
score on 1-MAE, the accuracy measure the research industry uses.
Up to 92%
on NDAM, a stricter measure we hold ourselves to, where GPT-5 scores 65%.
50k+
evaluations against real survey responses, across 155 countries.
The problem with how accuracy is measured today
Ask the same person the same question twice in one survey and they agree with themselves around 94% of the time. People misread, rush and change their minds, so there is no perfect ceiling even in the real world.
What would perfect even look like?
There is no 100% accuracy. Because humans are inconsistent, 94% is the realistic human ceiling for any research, including ours.
How are you measuring?
We report both the industry standard 1-MAE score, and the NDAM score to measure our audience against. The 94% is the NDAM score on the traditional research ceiling, which is the stricter measure out of the two.
So where do you land?
We land at 92% against the 94% traditional research benchmark on NDAM.
Why should I trust your measurement?
Because we publish the lowernumber. On the industry-standard measure, 1-MAE, we score up to 96%.
1-MAE gets more generous as questions gain answer options, so it can return a more flattering score. We quote NDAM instead.
How we measure
Two measures: the industry standard and our stricter one
Definitions
1-MAE is the industry benchmark measure, and we report it so you can compare us like-for-like.
It measures how closely we match the average answer across all respondents.
NDAM is our stricter way of measuring accuracy. It measures how closely we predict the distribution of all responses, so we can show you the full spread of answers respondents give. It scores the shape of the whole distribution rather than whether we got the top-line answer right, and it holds the same meaning whether a question has two options or ten.
Why this matters
A high NDAM score means we can show you who your outliers are, who dissents, and where real opinion is split, instead of one garbled average voice. Your customers aren’t the same, so their responses shouldn’t be either.
The industry standard
Up to 96%
1-MAE (one minus mean absolute error)
Our stricter standard
92%
NDAM (normalised distribution accuracy measure)
How we compare
How we stack up against traditional research and available LLMs
Against traditional research, our synthetic audiences achieve a 92% NDAM, making it almost as reliable as the real thing. An available frontier model such as GPT-5 scores 65% on the same measure.
We compare against available LLMs so you can see how we do against the other models. We’ve found that LLMs are good at producing a largely correct-sounding average answer, but they don’t do well predicting what each individual in your dataset would say.
How we test
How we know our audiences are reliable
Prof. Michael Muthukrishna
Chief Science Adviser, Electric Twin
How we prove it
We test our accuracy on your data, before you take our word for anything
Every time we onboard a new customer, we take their real-world data, build a synthetic audience and run evaluations to measure how closely our synthetic audience predicts their real survey results.
How it works
The Times result
Case study
The Times tested our synthetic audience against their 642,000 subscriber-base, and got a 92% match (NDAM)
“We went from rationing research to running it on demand. The same team, ten times the output, and we could prove it matched our real audience.”

Chris Courtney-Smith
Director of Data Operations, The Times
Papers and research
We don’t just claim it, we’ve proved it
Read our published research, or download our white paper to learn more about the science
Published research

Where You Ask Matters: Platform Effects of Online Surveys and AI-Based Human Subject Simulation
March 2026
Where it does not work
What we would not use Electric Twin for
Audiences with too little data
Accuracy depends on the depth of real audience data behind the model. For very thin or newly formed audiences, we may recommend primary research first.
Claims that require human respondents
Regulated claims, clinical studies and legally mandated consumer testing need real people on record. Synthetic audiences do not replace those, and we will not tell you they do.
A replacement for your research programme
Electric Twin works alongside surveys, not instead of them. Traditional research remains the yardstick we measure ourselves against, and periodic real-world validation is how the models stay calibrated.
FAQs
Frequently asked questions
1-MAE is the industry standard, so it lets you compare us like-for-like against any survey method. NDAM is the stricter standard we hold ourselves to. It allows you to compare how accurate we are in predicting the range of responses instead of just one average answer.See how we explain the difference
Share it with your team
