How to use ChatGPT for market research, and where it falls short

How to use ChatGPT for market research, and where it falls short

ChatGPT is genuinely useful for four research jobs and unreliable at a fifth: answering as your customers. Here are the prompts that work, the checks that keep it honest, and the number behind the gap

A plausible answer is the most dangerous kind. In Everything Is Obvious, the sociologist Duncan Watts makes the case that common sense hands us a plausible-sounding story about why things happen, and the story feels like understanding even when it explains nothing. Ask a general-purpose model a research question and you get the same effect in seconds. The answer comes back fluent, reasonable and well organised, and you are left to work out whether any of it is true.

For some research jobs that is fine, because the model is working through material you gave it. For one job it is not, and that job is answering as your customers. There is a number on the gap. On the measure that scores how well a model predicts a real spread of answers, GPT-5 reaches 65%.

What is ChatGPT actually good at in market research?

The four jobs it does well share one feature. In each, the model is working through material you handed it rather than inventing opinions it has no basis for.

Summarising desk research. Paste in analyst reports, internal decks and interview notes, and ask for market context, the competitive set, the recurring objections and the hypotheses worth testing. The model is compressing text you gave it, which is the task current models are best at.

Drafting questionnaires and discussion guides. Ask for a draft, then ask the model to attack its own draft for leading questions, double-barrelled questions and assumed behaviours. The second pass is where the value is. Most people stop after the first and ship a worse instrument than they would have written alone.

Coding open-ended responses. Give it a coding frame and four hundred free-text answers, and ask it to assign each answer to a code and flag the ones that fit nothing. This is classification against a structure you defined, and a good use.

Mapping a competitive set. Ask for direct competitors, adjacent categories and the budget line your product really competes with. Then check every name, because the model invents companies with the same confidence it names the real ones.

What are the best ChatGPT prompts for market research?

These four work. One instruction improves all of them, so start there.

Give the model a role and a stated limit in the same breath. Something like: "You are a research director. If a claim isn't in the material I've pasted, say you don't know rather than filling the gap." Models fill gaps by default. Telling one not to reduces how often it does, though it will not stop it entirely.

1. Desk research synthesis.

Here are my notes from six stakeholder interviews and two analyst reports. Produce a market context summary, the five buyer objections that appear most often, and three hypotheses I could test. Cite which document each point came from. Where the documents disagree, say so rather than resolving it.

The instruction not to resolve disagreement is the important half. Left alone, a model smooths two contradictory sources into one confident paragraph, and the contradiction was the finding.

2. Questionnaire drafting and self-critique.

Draft twelve questions to test whether our subscribers would pay more for an ad-free tier. Then review your own draft and list every leading question, every double-barrelled question, and every question that assumes a behaviour I haven't established.

3. Coding open-ends.

Here's a coding frame with eight codes and four hundred free-text answers. Assign each answer to exactly one code. Return a table of answer, code and confidence. List every answer that fits no code separately rather than forcing it.

4. Competitive mapping with a verification pass.

List direct competitors, adjacent categories, and the internal budget line our product competes with, for a mid-market insight platform sold to UK enterprises. For each named company, state what you're confident about and what I need to verify independently.

How do you stop ChatGPT inventing things?

None of these is optional if the output is going anywhere near a decision.

Give it the source rather than asking from memory. Hallucination rises sharply when a model is recalling rather than reading. Paste the material.

Hand-check a sample. For coded open-ends, check thirty by hand before trusting the other three hundred and seventy. If the sample is clean, the batch probably is. If two of thirty are wrong, the frame is wrong.

Ask for the confidence and the gap. A model asked to flag what it is unsure about will flag some of it. One that is not asked will volunteer none.

Never let a chat transcript be the evidence. This is what catches teams out in front of a sceptical head of insight. There is nothing to check a transcript against, so it does not survive the first challenge.

So how far off is it when you ask it to be your customer? A validated result survives the challenge because someone can check it. When The Times scaled its research, the part that survived scrutiny was the part it could check.

"We went from rationing research to running it on demand. The same team, ten times the output, and we could prove it matched our real audience."

Chris Courtney-Smith, Director of Data Operations, The Times

What does ChatGPT get wrong, and by how much?

It answers as everyone at once. A frontier model trains on the widest data it can reach, which makes it good at producing a plausible average and gives it no way to anchor to the specific group you care about. Ask what your subscribers think, and it reports a composite of the internet with your question's framing on top.

The size of that problem is measurable. On NDAM, which scores how closely a model predicts the full distribution of real answers rather than just the average, GPT-5 reaches 65%. A synthetic audience built from real customer data and scored against held-out answers reaches up to 92% on the same measure. For reference, the human ceiling is around 94%: ask a real person the same question twice in one survey and they agree with themselves about 94% of the time.

The gap between 65% and 92% is largely about outliers. A plausible average gives you the headline the room had already guessed. It cannot tell you which segment objects, and the objecting segment is usually why a launch underperforms.

There is a second difference that gets less attention, and that is variance. We ran a head-to-head against ChatGPT on predicting survey response distributions using the Gallup World Poll, and won on average across every country and all twelve question categories, with much smaller variance in our predictions. Nate Silver devotes much of The Signal and the Noise to this failure mode. Before the 2008 crash, the credit-rating agencies stamped confident, precise-looking grades on mortgage securities that turned out to be close to worthless, and nothing on the face of a rating warned anyone which ones. A general-purpose model has a smaller version of the same problem. It is less accurate on average, and it gives you no signal about when it is badly wrong. An error you cannot see coming is worse than a larger one you can correct for.

Show Image NDAM accuracy: a grounded synthetic audience against a persona prompt in a frontier model. Sources: Electric Twin accuracy methodology, and Electric Twin: Platform Accuracy, August 2026.

One clarification, because it gets muddled. We build our audiences using large language models too, frontier models included. The difference is the survey data those audiences are built from, how they are built, and what happens to the output afterwards. We are not against language models. We are against asking an ungrounded one to stand in for your customer.

"We use AI for research" describes almost nothing. It is the market research equivalent of saying you use computers. What matters is whether the model is grounded in data about the specific people you are asking about, and whether anyone has ever scored its answers against reality.

When should you use ChatGPT, and when do you need a synthetic audience?

Ask one question: is the model working through your material, or predicting people it has never met?


Research job

Tool

Why

Summarising documents you already have

ChatGPT

Compressing supplied text. No prediction.

Drafting a questionnaire or discussion guide

ChatGPT

A first draft a researcher then edits. Cheap and fast.

Coding open-ended responses

ChatGPT, with a hand-checked sample

Classification against a frame you defined.

Mapping a competitive set

ChatGPT, then verify every name

Useful recall, unreliable precision.

Predicting how your customers will respond

Synthetic audience

65% against up to 92% on NDAM. Needs grounding in real answers.

Reading how one segment differs from the average

Synthetic audience

Requires respondent-level construction, not a composite voice.

Evidence a sceptical stakeholder will accept

Synthetic audience

A holdout score is checkable; a chat transcript is not.

Anything physical or regulated

Neither

Real respondents, on the record.

What is the difference between a synthetic audience and a persona prompt?

Grounding and scoring.

A persona prompt tells a general-purpose model to behave like a customer type. The model complies, fluently, and there is no way to tell whether the answer resembles anyone real. A synthetic audience is built from real survey or first-party data about specific people, and its answers are scored against real answers it was never shown.

The scoring is the part a prompt cannot replicate at any level of sophistication. We split a real dataset, build from one part, hold the other back, and report how closely the synthetic answers match, using two measures rather than one. The Times ran that protocol on a 10% holdout of its 642,000 subscribers and got a 92% match, against roughly 93% for the traditional methods it was already using.



Persona prompt in ChatGPT

Synthetic audience

Built from

A description you typed

Real survey or first-party customer data

Anchored to

Nothing specific

A named, real audience

Validated against

Nothing

Held-out real answers, never shown to the model

Reports outliers

No, one averaged voice

Yes, the full distribution

Score on NDAM

65% (GPT-5)

Up to 92%

Survives a sceptical stakeholder

Rarely

Where the holdout is on the record

The longer version is in why generic LLMs aren't synthetic audiences, and the mechanics are in what a synthetic audience is and how it works.

Does ChatGPT Deep Research change the answer?

For desk research, it helps. Deep Research runs a longer retrieval loop and returns a cited report, which makes it a better version of the first job on this page. Still check the citations, because a cited claim can misrepresent the source it points at.

For predicting your customers, it changes nothing. A longer search cannot fix the problem, because the people you are asking about are not in the public material it searches. Reading more of the internet does not put your subscribers in it.

Is Claude or Perplexity better than ChatGPT for market research?

They have different failure profiles, and the honest answer is that the choice matters less than what you ask any of them to do.

Perplexity is built around retrieval and citation, which suits desk research and makes verification faster. Claude handles long pasted documents well and tends to hold a coding frame consistently across a large batch, which suits open-end classification. ChatGPT has the widest tooling and the strongest automation ecosystem.

None of that changes the boundary. Every general-purpose model on that list answers as a composite of its training data when you ask it to be your customer. Swapping one for another moves audience-prediction accuracy by a small amount from a low base. The 65% figure belongs to ungrounded models, not to one vendor. So pick whichever handles your documents best, and do not ask any of them to be your customer. The head-to-head detail is in Electric Twin compared with LLMs.

What does a good AI-assisted research workflow look like, end to end?

The order does most of the work.

One, start from the decision. Write down the decision the research has to serve and what you will do differently depending on the answer. Research that cannot change an action rarely gets read.

Two, use a model to shape the question. Draft the instrument, then have the model attack its own draft. Fix the leading questions before anything goes out.

Three, run the question against a grounded audience. This is where a synthetic audience does the work a general-purpose model cannot, because the answer has to come from a model of the people you are actually asking about.

Four, read the distribution, not the average. Look for segments that disagree. Where two parts of one audience want opposite things, the average is a number nobody holds.

Five, probe the surprises. Ask a follow-up on anything that contradicts what the team believed. A grounded audience holds up and gives you the reason behind the answer.

Six, validate periodically against real people. Not every question, but on a cadence. Real respondents are the calibration layer that keeps the whole workflow honest. How synthetic and traditional research fit together is set out in what happens when you give the same survey to seven panels and one AI.

A general-purpose model belongs in steps two and five. It does not belong in steps three and four.

Where does each option not work?

ChatGPT will invent specifics. Competitor names, statistics and citations arrive with the same confidence as accurate ones. Anything factual needs checking against a source.

ChatGPT cannot be your evidence. A transcript will not survive a sceptical head of insight, because there is nothing to check it against.

Synthetic audiences need real data behind them. Thin or newly formed audiences need primary research first.

Neither reaches regulated or physical questions. Regulated claims, clinical studies and legally mandated testing need real people on record. Taste, packaging and in-person usability need a body in a room.

We publish those limits on the accuracy methodology page.

Frequently asked questions

Can ChatGPT replace market research? No. It is useful for summarising material you supply, drafting instruments and classifying responses you already have. Asked to predict how a specific group would answer, GPT-5 scores 65% on NDAM against up to 92% for a synthetic audience grounded in real data.

Is ChatGPT accurate enough for early-stage research? For structuring a question, yes. For answering it, 65% on NDAM is the number to weigh. Directional use is defensible as long as you say plainly that the output is a hypothesis, not evidence.

Can ChatGPT analyse open-ended survey responses? Reasonably well against a coding frame you define, provided you hand-check a sample. Ask it to flag answers that fit no code rather than forcing them.

What sample should you hand-check? Thirty is a sensible floor for a few hundred responses. If two or more of thirty are miscoded, the problem is usually the frame, not the model.

Do teams use both ChatGPT and a synthetic audience? Commonly, and in that order: ChatGPT to shape the question, a synthetic audience to answer it, a real survey periodically to re-validate.

What accuracy do generic AI tools reach on audience prediction? GPT-5 scores 65% on NDAM. Frontier models produce a plausible average and struggle to predict what each individual in a dataset would say, which is the part that changes decisions.

Stop guessing.

Start predicting

Stop guessing.

Start predicting

Stop guessing.

Start predicting