A synthetic panel's one defensible job is search-space reduction: screening a wide field of ideas down to a shortlist that real people then test. You are about to run one, 2,000 AI respondents scoring eight snack concepts, fast and cheap. First, say what the run is for.
You asked the simulation to settle a question. That is exactly where it fails. The average looks right while the spread of opinion collapses, and about a third of these synthetic relationships point the wrong way, with no way to tell which (Bisbee et al., 2024). Change the wording or the model version and the ranking moves. A number you cannot reproduce next week is not a decision.
Good. A screen produces candidates to test, not answers to report. The panel ranks the eight and cuts the bottom three, so your real fieldwork goes to the five survivors.
The screen ranked the concepts. Whether a real test keeps the same order, rank preservation, is the property the whole method depends on, and the one no published study has confirmed. A model can call an effect's direction (Ashokkumar et al., 2026); the shortlist it cannot yet be trusted to. Here is how one real test scored the same eight:
The synthetic shortlist dropped two of the real top four, Oat cluster and Berry beet, and kept its own second favourite, Chili mango, that real people put seventh. On a real decision you would never have tested the winners it cut.
Three weeks later. A VP wants one line for the board deck: Miso caramel preferred, 62 percent. The word directional has quietly dropped off it. Do you paste it in?
That is how almost every misuse happens. A filter result, with the disagreement flattened out of it and its ranking never checked, is now a finding a decision rests on. The real customers who were never asked still have to show up. Same number, wrong slot.
The simulation widened what you looked at and ranked it for you. Whether that ranking matches real people is still the open question, so real people decide what is true. A directional finding is a reason to test something, never a reason to skip the test.
Illustrative panel outputs, not recorded model output; a Webflow embed cannot call a model. The synthetic and real-test rankings and the 62 percent are invented to show the mechanic; no published study has shown that a synthetic screen preserves the ranking of a real test, and that a model can call an effect's direction but not its magnitude is Ashokkumar et al. (2026). The one-third-point-the-wrong-way figure is measured (Bisbee et al., 2024). Updated 2026-08-26.