Message testing with synthetic audiences: what to test, and what not to trust

Message testing with synthetic audiences: what to test, and what not to trust

Put every route in front of a model of your real customers and read the response by segment, in about a minute. Here is what to test, how to ask, and what not to trust

Some years ago the sociologist Duncan Watts and his colleagues built a website that let about 14,000 people listen to and download songs by bands they had never heard of. One group saw only the song titles. Another group also saw how many times each song had already been downloaded, and that group was split into 8 separate copies of the site, each sealed off from the others, each starting from zero downloads. The same song rose to the top in one copy and sank in another. A track called "Lockdown" by 52 Metro led one world and trailed in the next. Gathering more information about the songs did nothing to make the winner easier to call.

The finding travels well beyond music. The qualities you can weigh at your desk, the wording, the logic, the taste of the room, explain less about how a message performs than the audience receiving it does. And the audience is the part you are guessing about.

Most message decisions still get made in a room. Three routes go in, and the one that comes out tends to be whichever the most senior person preferred. A synthetic audience changes what that meeting has to work with. You can put every route, tagline or piece of creative in front of a model of your real customers and read how each segment responds in about a minute, before you commit a penny. There is no fieldwork cost, so the eleventh option costs the same to test as the first. The eleventh is sometimes the one that wins.

What can you test with a synthetic audience?

Anything you can put into words or show as a picture, before you spend on it:

  • The core proposition, and which of three framings of it lands

  • A campaign tagline or hook

  • Proof points, and which reason to believe carries weight with which segment

  • A landing page headline and subhead

  • Social and search ad copy

  • An email subject line and opening angle

  • The category framing you use to explain yourself

  • A competitor comparison

  • The call to action

  • A product, format or podcast name

  • Creative and image routes, through creative testing

  • A price point, through pricing testing

Most of that list never gets tested. A subject line, or a shortlist of three taglines, is too small to justify commissioned fieldwork, so it gets settled by opinion. That gap is wider than the big set-piece studies, and it is where most day-to-day decisions actually live.

How does a message test with a synthetic audience run?

Four steps: build the audience, load every option, ask well, read by segment. Teams tend to be strong on the first two. The third is where tests go wrong.

One, build the audience from your data. Real survey or first-party customer data about the people you are targeting. The stages are set out in how synthetic audiences work.

Two, load every option, including the ones you expect to lose. Adding a fifth costs nothing, and the fifth sometimes wins.

Three, ask a question that forces a decision. "Which do you prefer?" returns a ranking you cannot act on. Ask what the message made the reader expect, what they would do next, and what would hold them back.

Four, read it by segment. The all-audience average is usually the least useful number in the output. The distance between two segments is where the decision actually sits.

What does a good message-test question look like?

Specific about the decision, and open enough to surface an objection you had not thought of.


Instead of asking

Ask

Because

Which of these three taglines do you prefer?

After reading this tagline, what do you think this product does?

You learn whether the message reads the way you intended.

Is this message clear?

What would stop you acting on this?

It names the objection rather than confirming the copy.

Do you like this offer?

What would you expect this to cost, and what would make it worth that?

It turns a "like" into a price expectation you can test.

Rate this ad out of ten

Who is this ad for, and is that you?

It catches a message aimed at the wrong segment.

Does this sound credible?

Which part of this would you want proof of?

It names the proof point that is missing.

Would you buy this?

What would you do in the week after seeing this?

A described next action tracks behaviour better than stated intent.

The through-line: questions that invite approval get approval, from real respondents and synthetic ones alike, and approval rarely tells you what to change. Ask instead what the reader understood, expected, and would do.

What changes when you read results by segment?

Reading by segment catches the messages that test well on average and fail with the people you are trying to reach, and the ones that do the reverse. Darrell Huff made the underlying point seventy years ago in How to Lie with Statistics: the word "average" quietly hides which average you mean, and a single figure can sit neatly between two groups who want opposite things.

The Times marketing team saw this when they tested four names for a new business podcast. The winner, The Business, did well with a general reader panel and did markedly better with the specific group the show was made for. Average the two together and that edge disappears, and the choice slides back to whichever name the room liked.

Product work showed the same shape. The paper's long-standing subscribers were wary of AI features. The readers it hoped to attract wanted exactly those features. A single averaged answer would have pointed the roadmap at no one in particular, and it would have set the roadmap anyway. Broad audience generalisations fail for the same reason, which is the subject of why Gen Z generalisations are bad for business. The detail on the Times work is in the Times case study.

What outputs should you expect from a test?

If a tool gives you only the first of these, you cannot act on it.

A ranking, by segment. Which route won overall, and, separately, which won with each segment you care about. The overall winner and the segment winner are often different, and that difference is frequently the finding.

The reasoning behind each answer, in the audience's own words. This is where a popular but wrong message shows itself: it scores well because it is being misread.

The objections, named. What would stop each segment acting. These map straight onto the proof points you are missing, and they are often more useful than the ranking itself.

The spread. Where opinion is genuinely divided, the output should say so rather than returning a midpoint no one holds. A single confident number for a split audience hides the thing you needed to see.

Feed the named objections back in as new variants. Two or three rounds of that, at seconds each, gets you further than one traditional test.

How does this compare with running a real message test?

A real message test is the benchmark, and it is slow enough that most messages never get one. Recruitment and fieldwork push a campaign test weeks out, past most approval deadlines, so the test either gets skipped or gets run on the single route already chosen.

Synthetic testing changes the order. Test the routes while they are still routes, using message testing, then take the survivor to a real test if the spend justifies it.

"We went from rationing research to running it on demand. The same team, ten times the output, and we could prove it matched our real audience."

Chris Courtney-Smith, Director of Data Operations, The Times



Traditional message test

Synthetic audience test

Time to result

Weeks

In seconds

Realistic number of variants

Two or three

As many as you can read

Cost per extra variant

Fieldwork cost

None

Confidentiality

Exposed to a panel

Stays internal

Segment-level read

Needs sample size in every cell

Available from the same run

Accuracy on NDAM

Around 94%, the human ceiling

Up to 92%

Best used for

Signing off the chosen route

Choosing which route to sign off

How accurate is a synthetic audience on creative?

We checked reactions to creative against a human benchmark of four thousand participants, across eighteen images and six videos of up to three minutes each. The test data was sealed off: none of the opinions people held about the assets were used to build the audiences, so the model was reacting to the creative itself rather than to a hint about what anyone thought of it. The predictions tracked the real reactions closely, including on which of two assets won a head-to-head, the shape most creative decisions take.

Across the platform we report up to 96% on 1-MAE and up to 92% on NDAM, both measured against real answers held back from the model. The method is set out in how accurate synthetic audiences are.

For a creative decision, the honest reading is this. It is strong enough to cut a shortlist to the one or two routes worth backing, and not strong enough to be the last word before the media spend goes out. Narrow with it, sign off with judgement, and use a real test where the budget warrants one.

What should you not test this way?

Three cases, and the third is the one that catches people out.

A regulated claim. Wording that carries legal exposure needs real respondents on record.

A final gate on major spend. At significant budget, a synthetic test narrows the shortlist and a real test signs it off. Reading the synthetic result as the sign-off is the most common way teams overreach.

Anything sensory. How a pack feels in the hand, how sound design fills a room, how something tastes. Those depend on a body in the room, not a reported opinion.

The standing limits apply too. Audiences with too little real data behind them need primary research first, and synthetic testing runs alongside a survey programme rather than replacing it. There is more on how the two fit together in what happens when you give the same survey to seven panels and one AI.

Frequently asked questions

How many message variants can you test at once? There is no fieldwork cost per variant, so the limit is how many you can read, not how many you can afford. Teams usually start with the three they were about to argue over, then add the two they had already discarded.

Can you test images and video as well as copy? Yes, through creative testing, and both are validated against a four-thousand-participant human benchmark across eighteen images and six videos of up to three minutes.

How do you know a message-test result is right? A portion of your real data is held back, the audience answers questions drawn from it, and we score those answers against the real responses. We report up to 96% on 1-MAE and up to 92% on NDAM against those held-out answers.

Does this replace A/B testing in market? No. An in-market test measures behaviour against real spend. A synthetic test decides what deserves to go into an in-market test. It sits earlier in the sequence and answers a different question.

What if the synthetic result contradicts what the team believes? Ask a follow-up. An audience grounded in your real data holds up under questioning and gives you the reasoning behind its answer, which is usually what shifts a view in the room.

Stop guessing.

Start predicting

Stop guessing.

Start predicting

Stop guessing.

Start predicting