Synthetic Respondents and Research Simulation

The most radical claim in market research right now is that useful answers can come from simulated people. This chapter takes that claim seriously in both directions — what synthetic respondents have actually demonstrated, and the quiet ways they flatten, parrot and mislead.

M2 / 11 published articles

What if you could interview someone who does not exist?

Somewhere in the last two years, a strange new practice entered the research industry. Instead of recruiting real people and asking them questions, you describe a person to a language model — age, income, attitudes, geography — and ask your question as if that person were sitting across the table. The model answers in character. Do it a thousand times with a thousand profiles and you have something that looks, statistically, like a survey panel. No recruitment, no fieldwork, no incentive fraud, no three-week wait. Results by morning.

This is silicon sampling, and it works better than sceptics expected and worse than vendors claim. On broad attitudinal questions the model's answers track real survey distributions surprisingly well. On anything that depends on lived experience, local knowledge, or genuine surprise, the answers converge toward a flattened, agreeable, suspiciously articulate average. The model does not know what it is like to be a person; it knows what text about that kind of person looks like. That distinction is the fault line running through this entire chapter.

The chapter's scope is wider than one technique. Digital twins, agent-based markets, synthetic pretesting, simulated rare audiences, LLM-augmented conjoint — these are all versions of the same bet: that a computational model of a person can stand in, at least partially, for the person. Some of those bets are paying off in specific, bounded uses. Others are being sold as something they are not. The job of this chapter is to help you tell which is which.

Why this matters to you

If you buy research, you are going to be offered synthetic alternatives. Some vendors are already shipping them without the label. That makes this chapter urgent in a practical way: you need to know what a synthetic result can actually deliver, what questions it can answer well, where it fails silently, and what to demand before you trust it. You also need a mental model for the more sophisticated uses — calibration, hybrid designs, pretesting — where the synthetic component adds genuine value precisely because it is not asked to replace the real respondent but to supplement one. The difference between those two uses is the difference between a tool and a trick.

What you'll find inside

The pieces here map the synthetic turn from its boldest claims to its hardest limits. You will start with silicon sampling itself and what the benchmarks actually show, then meet digital twins — persistent simulated individuals that carry a person's profile across time — and the fidelity problem they have not solved. You will confront the flattening, the systematic ways synthetic populations lose the tails, the surprises, the contradictions that make real data valuable. Then the chapter builds toward the uses that work: calibration and hybrid designs that splice synthetic and human data, pretesting that runs a rehearsal audience before a real one, agent-based markets that simulate competitive dynamics, and the targeted simulation of rare populations no panel can reach. You will see what happens when conjoint analysis meets a language model, and how to audit a vendor who says their synthetic product is ready. The chapter closes with two framing pieces: what simulation is actually for when you stop expecting it to replace truth, and why the word "synthetic" itself now means two completely different things that the industry keeps conflating.

The honest note

The characteristic caveat for this chapter is inherited straight from the behavioural science that preceded it: if people cannot reliably tell you why they do what they do — and decades of evidence say they cannot — then a model of people cannot either. A synthetic respondent is a model of what a person says, not of what a person is. That is useful for some purposes and actively misleading for others, and the line between the two is the thing the industry most urgently needs to learn to draw. The vendors pitching "unlimited sample, instant turnaround" are not lying about the capability; they are lying about the trade-off. This chapter exists so you can see the trade-off clearly.

The eleven pieces in this chapter

  1. "Ask the model what the public thinks" — silicon sampling
  2. "A copy of the customer, four-fifths faithful" — digital twins
  3. "Where synthetics fail" — the flattening
  4. "Splice, don't substitute" — calibration and hybrids
  5. "The rehearsal audience" — synthetic pretesting
  6. "A thousand shoppers in a bottle" — generative agent markets
  7. "Synthetic surgeons and CFOs" — simulating the unreachable
  8. "Preference measurement, part-machine" — conjoint meets the model
  9. "The demo always works" — auditing a synthetic vendor
  10. "A hypothesis engine, not a truth machine" — what simulation is for
  11. "Privacy copies vs pretend people" — two meanings of 'synthetic'

Read the articles

  1. M2-01

    Synthetic Survey Respondents: Ask the Model What the Public Thinks

  2. M2-02

    Respondent Digital Twins: The Twin Knows Your Type, Not You

  3. M2-03

    Synthetic Sample Validation: The Mean Matches. That Is the Trap.

  4. M2-04

    Survey Sample Augmentation: Borrow the Answers You're Missing, Up to a Point

  5. M2-05

    Synthetic Data for Survey Pretesting: Fake the Survey, Not the Findings

  6. M2-06

    AI Agent Market Simulation: Run the Same Market Twice, Get Two Worlds

  7. M2-07

    Synthetic Expert Interviews: Interview the Expert You Could Never Book

  8. M2-08

    Conjoint Analysis with Fewer Respondents: Cut the Panel, Keep the Referee

  9. M2-09

    Synthetic Research Vendor Audits: The Demo Always Works

  10. M2-10

    Synthetic Respondents for Hypothesis Generation: A Hypothesis Engine, Not a Truth Machine

  11. M2-11

    Evaluating Synthetic Data Claims: One Word Now Covers Two Very Different Products

Let’s talk

Invisible forces shape your world — until you hire Latenta®

Contact