Two research firms hand you the same conjoint study. Each cut its panel of real people to a
fraction of the usual size and let a language model fill in the missing preferences, so each cost
far less to run. Both reports reach the same verdict: your new bundle should win. Their preference
weights, the bars below, look the same too. Which result can you trust?
Firm A
Recommendation
Price weight
Warranty weight
Design weight
Launch the bundle.
Firm B
Recommendation
Price weight
Warranty weight
Design weight
Launch the bundle.
Pick a result first. The referee is a fresh sample of real people the models never saw.
Illustrative worked example, not measured data. The held-out-sample
reduction it dramatises is real: a preprint (Wang, Zhang and Zhang, 2024) recorded a 24.9 to 79.8
percent cut in real human choices across two categories, checked against real people the model
never saw.