Skip to main content

Accuracy

Up to 96% accurate against real people’s answers

Up to 96% accurate against real people’s answers

Up to 96% accurate against real people’s answers

As accurate as traditional research, with published proof so that you can check the claims yourself

Up to 96%

score on 1-MAE, the accuracy measure the research industry uses.

Up to 92%

on NDAM, a stricter measure we hold ourselves to, where GPT-5 scores 65%.

50k+

evaluations against real survey responses, across 155 countries.

What it means to us

Accuracy is central to everything we do

Up to
Up to 96%
accurate (1-MAE) against real people’s answers

At Electric Twin, accuracy is at the core of how we view our work. Our synthetic audiences are not just large language models with opinions. We ground all our answers in data and run evaluations consistently, so that your audience’s answers track what your real customers would say.

01 · Measured

Against real answers

Every score comes from comparing answers from our synthetic audience against those of real humans.

02 · Published

Two numbers, not one

We report the industry measure and a stricter one for transparency.

03 · Within limits

We tell you what we can’t do

There are questions we will tell you to take to human respondents instead.

The problem with how accuracy is measured today

No research method gets it exactly right, including the ones you already use

No research method gets it exactly right, including the ones you already use

Ask the same person the same question twice in one survey and they agree with themselves around 94% of the time. People misread, rush and change their minds, so there is no perfect ceiling even in the real world.

What would perfect even look like?

There is no 100% accuracy. Because humans are inconsistent, 94% is the realistic human ceiling for any research, including ours.

How are you measuring?

We report both the industry standard 1-MAE score, and the NDAM score to measure our audience against. The 94% is the NDAM score on the traditional research ceiling, which is the stricter measure out of the two.

So where do you land?

We land at 92% against the 94% traditional research benchmark on NDAM.

Why should I trust your measurement?

Because we publish the lowernumber. On the industry-standard measure, 1-MAE, we score up to 96%.

1-MAE gets more generous as questions gain answer options, so it can return a more flattering score. We quote NDAM instead.

How we measure

Two measures: the industry standard and our stricter one

Definitions
1-MAE is the industry benchmark measure, and we report it so you can compare us like-for-like.

It measures how closely we match the average answer across all respondents.

NDAM is our stricter way of measuring accuracy. It measures how closely we predict the distribution of all responses, so we can show you the full spread of answers respondents give. It scores the shape of the whole distribution rather than whether we got the top-line answer right, and it holds the same meaning whether a question has two options or ten.

Why this matters
A high NDAM score means we can show you who your outliers are, who dissents, and where real opinion is split, instead of one garbled average voice. Your customers aren’t the same, so their responses shouldn’t be either.

The industry standard

Up to 96%

1-MAE (one minus mean absolute error)

Our stricter standard

92%

NDAM (normalised distribution accuracy measure)

How we compare

How we stack up against traditional research and available LLMs

Against traditional research, our synthetic audiences achieve a 92% NDAM, making it almost as reliable as the real thing. An available frontier model such as GPT-5 scores 65% on the same measure.

We compare against available LLMs so you can see how we do against the other models. We’ve found that LLMs are good at producing a largely correct-sounding average answer, but they don’t do well predicting what each individual in your dataset would say.

How we compare on our NDAM standard
0%
Electric Twin
NDAM score
0%
Traditional research
NDAM score
0%
GPT-5
NDAM score

How we test

How we know our audiences are reliable

Prof. Michael Muthukrishna

Chief Science Adviser, Electric Twin

How we prove it

We test our accuracy on your data, before you take our word for anything

Every time we onboard a new customer, we take their real-world data, build a synthetic audience and run evaluations to measure how closely our synthetic audience predicts their real survey results.

How it works

01
Split a dataset into two
We take a real survey dataset and split it into two parts.
02
Build from one part
The first section is used to build the synthetic audience.
03
Hide the other part
The second section is ‘held out’ and never shown to them.
04
Test and compare
We test the audience on questions they never see, then score how closely their answers match what real humans said.

The Times result

Case study

The Times tested our synthetic audience against their 642,000 subscriber-base, and got a 92% match (NDAM)

“We went from rationing research to running it on demand. The same team, ten times the output, and we could prove it matched our real audience.”

Chris Courtney-Smith, Director of Data Operations, The Times

Chris Courtney-Smith

Director of Data Operations, The Times

Papers and research

We don’t just claim it, we’ve proved it

Read our published research, or download our white paper to learn more about the science

Published research

Electric Twin published research documents

Where You Ask Matters: Platform Effects of Online Surveys and AI-Based Human Subject Simulation

March 2026

White paper

Electric Twin Platform Accuracy white paper

Electric Twin: Platform Accuracy

August 2026

Where it does not work

What we would not use Electric Twin for

Audiences with too little data

Accuracy depends on the depth of real audience data behind the model. For very thin or newly formed audiences, we may recommend primary research first.

Claims that require human respondents

Regulated claims, clinical studies and legally mandated consumer testing need real people on record. Synthetic audiences do not replace those, and we will not tell you they do.

A replacement for your research programme

Electric Twin works alongside surveys, not instead of them. Traditional research remains the yardstick we measure ourselves against, and periodic real-world validation is how the models stay calibrated.

FAQs

Frequently asked questions

Why report two accuracy measures?

1-MAE is the industry standard, so it lets you compare us like-for-like against any survey method. NDAM is the stricter standard we hold ourselves to. It allows you to compare how accurate we are in predicting the range of responses instead of just one average answer.See how we explain the difference

How does Electric Twin prevent data leakage in accuracy testing?
Is accuracy improving over time?
How does this compare to a real survey’s margin of error?

Share it with your team

Send the white paper to the biggest
sceptic on your team