Conjoint Analysis with Fewer Respondents: Cut the Panel, Keep the Referee

Article M2-08

A language model can now feed extra preference data into a conjoint study and cut the number of real people you need by a quarter to four-fifths. The saving holds only while real respondents still get the last word, and no AI conjoint vendor sells it yet.

In brief

Conjoint analysis is the survey method that works out what people value by making them choose between product options, then reading how much each feature moved the choice. It is powerful and expensive, because a reliable result needs a lot of people making a lot of choices. The post-2022 claim is that a language model can reduce that cost. Instead of the model pretending to be a respondent, it feeds extra, model-generated preference data into the estimation, and a smaller group of real people anchors the result. On two demonstrations this cut the number of human choices needed by somewhere between a quarter and four-fifths. The evidence is limited and comes mostly from one side. The reduction rests on a single preprint and its follow-ups, none yet reviewed, while the peer-reviewed papers are the sceptical ones. The saving occurs only when a held-out sample of real people checks the answer. No vendor sells this, and the largest vendors argue against it.

What the theory says

The theory

Conjoint analysis is the survey method that works out what people value by observing them making choices. Instead of asking someone how important price or battery life is, which they cannot answer reliably, you show them a series of product options, each with different features, and ask which they would pick. From enough of those choices you can estimate how much each feature and each level influenced the decision. Those weights are called part-worths, and they are what a conjoint study provides: a number for how much a longer warranty is worth against a lower price, or which of five features a buyer would give up last.

The method is trusted, and it is expensive. A reliable result needs many people, each making many choices, because the estimate has to hold not just on average but across the different kinds of buyer in a market. Researchers have spent decades reducing that cost without losing the result. Adaptive designs ask each respondent the questions that reveal the most, given what they have already said. Efficient designs choose the set of options that get the most information from the fewest choices (Berni, Nikiforova and Pinelli, 2024). Hierarchical models let one respondent's answers inform the estimate for a similar respondent, so the sample is more efficient than its size suggests. The burden problem is not new, and the classic answers to it are real. Any new claim has to be measured against those, not against a naive full-length survey.

The post-2022 claim is that a language model can do part of the work. The version that gets the most attention, and the one this article is not about, is the model as a stand-in respondent: you ask the model to answer the conjoint as if it were a person, and skip the people. That idea has its own article at [M2-01], and the evidence for it is mixed at best. The move that is more interesting, and better evidenced, keeps real people in the study and uses the model in a different role. It comes in two forms.

The first form is used in the estimation. A model generates extra preference data, and that synthetic data is fused with a smaller sample of real human choices inside the statistical model that produces the part-worths. The real sample anchors the result and the model data supplements it, so you need fewer real choice sets for the same precision. This is the same splice-don't-substitute logic that [M2-04] sets out for thin survey cells in general, applied here inside the conjoint estimator. The primary source is a single preprint by Wang, Zhang and Zhang (2024), who call it a data-augmentation approach and are explicit that the model data is a complement and not a substitute. If used directly as a replacement for real answers, it can make bias worse, and the value comes only from combining it with real data in an estimator that carefully weighs each source. On two demonstrations this cut the number of human choices needed by between a quarter and four-fifths.

The second form is used earlier, in the design. Before a conjoint study runs, someone has to decide which attributes and levels to test, and a model can suggest them or generate the material respondents see. Brauner (2026) built an open-source platform that does exactly this, with the model proposing attributes and helping generate the stimuli. A related line uses model-based priors to choose the choice sets themselves, starting the design with the model's expectations (Eggers and Vriens, 2026). In both cases the model is a drafting tool, and a person still decides. A separate line of work improves the conjoint estimator with neural representation learning rather than a language model, which is closer to the modelling engines in [M4-09] than to anything here (Zhang et al., 2025).

The key element is the validation check. In every serious version of this, the answer is checked against a held-out sample of real people: real human choices the estimator never saw, used to test whether the augmented part-worths actually predict what people do (Wang, Zhang and Zhang, 2024). That held-out check is the principle covered in [M0-03], and it is the essential part. Without it, model-generated preference data is just an unverified claim about what people want. The reason direct model elicitation is not trusted on its own is that it fails this test. Goli and Singh (2024), in the one peer-reviewed result on direct model elicitation, found that asking a model for preferences directly is unreliable, which is the gap between stated and actual behavior in [M0-07], now applied to machines. The estimation-loop method survives only because it never asks the model to be right on its own. It uses a smaller group of real people to validate it.

Controversies

The main number needs careful attention. The claim that AI cuts a conjoint sample by roughly 25% to 80% comes from one place. The range in Wang, Zhang and Zhang (2024) runs from 24.9% to 79.8%, measured on two product categories, COVID-19 vaccines and sports cars. It is a held-out result, which is the good kind, but it is a single unreviewed preprint on two categories, not a span observed across many studies. The wide range is a warning sign. A saving that can be anywhere from a quarter to four-fifths depending on the category is not a fixed discount you can budget for in advance. Ye and Yoganarasimhan (2026) address the same question from a different angle, asking how much real human data you still need for the correction to hold, which balances the idea that 80% is guaranteed.

The second disagreement is about who is actually doing this. One might assume that if the saving were real, major conjoint vendors would already sell it. They are not, and the two loudest have argued publicly against the broader idea of synthetic survey data. Sawtooth Software (2026) published a position rejecting synthetic and model-generated data as survey data, and Conjointly (2026) documents its own AI features as survey-authoring and reporting aids while opposing synthetic respondents. What the market sells as AI for conjoint is help writing the survey and summarising the output, not a model used in the estimation. So the estimation-loop method is a research literature that the incumbents are pushing back on, not a product you can buy today.

The third open argument is more technical, and it matters most for anyone who uses conjoint properly. A conjoint study is usually valued for its individual-level detail: not just what the average buyer wants, but how a price-sensitive segment differs from a segment that wants features. The demonstrated augmentation result is strongest at the aggregate, average level. Whether it preserves that individual-level spread, or narrows everyone toward the middle, is not settled. The one strong individual-level result in this area, where model-based agents reproduced a specific person's choices at a high rate (Xuan, Hwang and Lee, 2026), sits on the model-as-respondent side of the line, not the estimation-loop side, so it does not close the gap. Kinzinger and Hartmann (2026) probe the same boundary and find that mimicking an individual respondent is hard. For now, the safe reading is that the saving is best evidenced for the average and unproven for the spread.

Limitations

The most important limitation is the state of the evidence itself, and it is unusual enough to state plainly. The main claim of the article, that a model can reduce your sample size, is based only on preprints: Wang, Zhang and Zhang (2024) and a small 2025 to 2026 cluster around it (Lu et al., 2026; Ye and Yoganarasimhan, 2026; Wang, Ye and Zhao, 2025; Eggers and Vriens, 2026). None has yet cleared peer review. The peer-reviewed papers in this area suggest the opposite. Goli and Singh (2024) find direct model elicitation unreliable, and Brand, Israeli and Ngwe (2024) find that model-derived willingness to pay is plausible in aggregate but unreliable in the specific details, which are the details a conjoint study is meant to provide. So the sceptical claims are the ones that have been reviewed, and the optimistic claim has not. That fact should be mentioned alongside the 80% figure whenever the figure is used.

The evidence for the design capability is even weaker. Brauner's (2026) platform generates attribute suggestions and stimulus material, but the model suggests and a human decides, and the study behind it involved 55 people looking at care robots. That is a working prototype, not evidence that a model can design a conjoint by itself. The suggestion itself has a less obvious risk. A model proposing the attributes to test may be introducing its own assumptions into the study before any person chooses, and as far as this research could establish, no one has yet checked whether generated attributes bias the design.

Then there is what is missing. A targeted search found no vendor selling a conjoint tool that puts a model inside the estimation loop and publishes a real-respondent holdout audit of the result. The cost saving has not been turned into a product, independently validated, or tested by an unbiased party. Every accuracy claim in this space is either an academic demonstration or a vendor's own report, and the two are different kinds of evidence.

And the citation record has few entries, which is easy to misinterpret. Across the searches behind this article, none of these preprints had yet drawn a published challenge or a published confirmation. In a field only months old, that silence does not mean agreement; it means nobody has checked the work yet, which is a reason to treat every number here as provisional.

Open questions

Three questions the field has not answered determine how much a model-augmented conjoint can be trusted.

Does the saving hold for a genuinely new product, or only where the model already knows the category? Every demonstration uses familiar categories, vaccines and sports cars, where a model has processed a large amount of relevant text, and none tested a thin category or respondents outside the mostly Western, English-language populations that dominate training data. A model is most useful where it already knows the answer and least trustworthy exactly where a conjoint study earns its fee, so a buyer cannot yet tell whether the discount holds for the studies they most need.

Can augmentation be shown to preserve the differences between buyers, not just the average? No clear held-out test at the segment level, comparing the price-sensitive buyer with the feature-driven one, has been published, so the issue remains unresolved. Until someone runs that test, the discount is safe to trust for aggregate questions but not for the segment-level detail that a conjoint study is usually purchased for.

Does the discount survive a change of model? Nobody has shown whether the saving holds across model versions or as models drift over time, and a discount that depends on which model you used last quarter is not something a study can rely on.

So what

A model generates extra preference data, a smaller sample of real people anchors it, and a held-out group of real people checks the result. Used that way, it can make a research budget go further. The same method can be misused in an obvious way. Preference data that a model generates can be steered. If you want a conjoint study to show that buyers love your bundle, or that voters back your policy mix, a model can be prompted toward data that says so, and the output will look like real research. The protection is the same thing that makes the honest version work: a held-out sample of real people who never saw the model's data and can contradict it. A conjoint study with no real holdout is not a cheaper study. The test to apply throughout is simple. Does a choice in the study serve the decision you are trying to make, or the answer you were hoping for? The rigorous version and the responsible version are identical.

For research practice

Treat model-generated data as a way to stretch a real sample, never as a way to remove it. The published method fuses synthetic data with real human choices and validates against a held-out real sample, so the real sample and the holdout are not the parts to cut. If a workflow uses a model to augment estimation, keep a block of real respondents the model's data never touches, and report the part-worths' accuracy against them. Report the evidence tier honestly too. The saving rests on preprints, the range is wide, and the peer-reviewed work is sceptical, so the defensible internal claim is that this is a promising method under test, not a settled way to run studies at half the cost.

For companies

The cost saving is real in principle and conditional in practice, and the pitch you will hear oversells it. When a vendor says its model matches human answers, ask what was matched and against what. The strong published result is a held-out reduction on two categories in an academic preprint, and the peer-reviewed finding on willingness to pay is that models get the aggregate roughly right and the detail wrong (Brand, Israeli and Ngwe, 2024). Willingness to pay in the detail is usually the number a pricing study is bought for. The safe way to capture the upside is to run the model-augmented estimate and a smaller conventional study on the same question once, compare them, and only trust the cheaper method on the categories where it already matched. Do not buy the version that removes real respondents entirely, because that is where the method's own authors say the bias creeps in.

For political parties

This is the audience where a manipulated result causes the most harm. Conjoint is used in politics to work out which bundle of positions a coalition will actually back, and model-augmented data makes it cheap to test many versions fast and in private. That same cheapness makes it easy to create the answer you wanted, because a model can be nudged toward preference data that flatters a chosen message, and the output still looks like a study. A survey engineered to agree with you is worthless as intelligence: you have paid to receive your own assumptions. So use the model to narrow a large field of message and policy combinations quickly, then test the survivors on real voters, including the ones a model impersonates worst. Those are the older, less formally educated, and second-language voters whose choices a model has the least basis to predict (Kinzinger and Hartmann, 2026).

For government and policy

Public bodies commissioning conjoint, for a policy trade-off or a service design, face the most serious disclosure issue. If model-generated data went into a result that guides a decision, that must be recorded. The 2025 revision of the ICC/ESOMAR International Code addresses synthetic data and synthetic respondents directly and adds a duty-of-care principle, and it is the reporting standard that complements the statistical standard (ICC/ESOMAR, 2025). The defensible standard for a public commission is to require three things: that any model-augmented estimate be validated against a real-respondent holdout, that the use of synthetic data be disclosed rather than hidden in a method note, and that the populations the model predicts worst be tested on real people. A number that affects public spending should not be based on data a model created and no real person verified.

How to use this

Before trusting a conjoint result that used a model, ask four questions. Did a sample of real people anchor the estimate, or did the model supply the preferences on its own? Was the result checked against a held-out group of real respondents the model did not see? Is the claim you are being sold the aggregate average, which the evidence supports better, or the individual-level detail, which it supports worse? And is the category one where a model has plausibly seen the relevant behaviour, or a genuinely new product where its confidence is least earned? If the answer to the first two is no, you do not have a cheaper conjoint study. You have a model's opinion that looks like a survey result but is not based on real respondents. The saving is worth having, but only if the checks are still in place.

Case studies

The held-out demonstration. The clearest evidence for the estimation-loop method is the pair of demonstrations in Wang, Zhang and Zhang (2024). The authors augmented conjoint estimates with model-generated preference data and checked the result against real human choice data held out from the estimation, on two categories: COVID-19 vaccine preferences and sports-car choices. The reduction in the number of human choices needed ran from 24.9% to 79.8% across the settings tested. This is an academic demonstration and not a client engagement, and it is a single preprint, so it should be read as the strongest existing proof of concept rather than as a shipped result.

The vendor that argued back. The most useful industry data point is a refusal. Sawtooth Software, one of the largest conjoint vendors, published a position against treating synthetic and model-generated data as real survey data (Sawtooth Software, 2026), and Conjointly (2026) describes its AI features as help with writing and reporting a survey rather than a model inside the estimation. Read together, they are evidence that the estimation-loop method has not become a product, and that the people best placed to sell it are, for now, arguing against the broader idea. That is a finding in its own right. The saving lives in the literature, not yet on the market.

References

Berni, R., Nikiforova, N.D. and Pinelli, P. (2024) 'An Optimal Design through a Compound Criterion for Integrating Extra Preference Information in a Choice Experiment: A Case Study on Moka Ground Coffee', Stats, 7(2), pp. 521–536. Available at: https://doi.org/10.3390/stats7020032 (Accessed: 18 August 2026).

Brand, J., Israeli, A. and Ngwe, D. (2024) 'Using GPT for Market Research', Proceedings of the 25th ACM Conference on Economics and Computation (EC '24), p. 613. Available at: https://doi.org/10.1145/3670865.3673479 (Accessed: 18 August 2026).

Brauner, P. (2026) 'From Prompts to Preferences: An Open-Source Platform for Generative AI-Enhanced Conjoint Analysis', arXiv preprint arXiv:2606.12972. Available at: https://doi.org/10.48550/arXiv.2606.12972 (Accessed: 18 August 2026).

Conjointly (2026) AI use cases and ROI in Conjointly. Available at: https://conjointly.com/blog/ai-use-cases-and-roi-in-conjointly/ (Accessed: 18 August 2026).

Eggers, F. and Vriens, M. (2026) 'Using LLM-Based Priors to Optimize Choice Designs in Conjoint Experiments', SSRN working paper 6719681. Available at: https://doi.org/10.2139/ssrn.6719681 (Accessed: 18 August 2026).

Goli, A. and Singh, A. (2024) 'Frontiers: Can Large Language Models Capture Human Preferences?', Marketing Science, 43(4), pp. 709–722. Available at: https://doi.org/10.1287/mksc.2023.0306 (Accessed: 18 August 2026).

ICC/ESOMAR (2025) ICC/ESOMAR International Code on Market, Opinion and Social Research and Data Analytics. Available at: https://esomar.org (Accessed: 18 August 2026).

Kinzinger, L. and Hartmann, J. (2026) 'Synthetic Personalities: How Well Can LLMs Mimic Individual Respondents Using Socio-Economic Micro-data', arXiv preprint arXiv:2606.04592. Available at: https://doi.org/10.48550/arXiv.2606.04592 (Accessed: 18 August 2026).

Lu, C., Wang, M., Zhang, D.J. and Zhang, H. (2026) 'Generative Augmented Inference of LLM-generated Data for Market Research: Theory and Empirical Evidence', arXiv preprint arXiv:2604.14575. Available at: https://doi.org/10.48550/arXiv.2604.14575 (Accessed: 18 August 2026).

Sawtooth Software (2026) Why synthetic data is not really data. Available at: https://sawtoothsoftware.com/resources/blog/posts/why-synthetic-data-is-not-really-data (Accessed: 18 August 2026).

Wang, L., Ye, Z. and Zhao, J. (2025) 'Efficient Inference Using Large Language Models with Limited Human Data: Fine-Tuning then Rectification', arXiv preprint arXiv:2511.19486. Available at: https://doi.org/10.48550/arXiv.2511.19486 (Accessed: 18 August 2026).

Wang, M., Zhang, D.J. and Zhang, H. (2024) 'Large Language Models for Market Research: A Data-augmentation Approach', arXiv preprint arXiv:2412.19363. Available at: https://doi.org/10.48550/arXiv.2412.19363 (Accessed: 18 August 2026).

Xuan, B., Hwang, J. and Lee, H. (2026) 'Your Reviews Replicate You: LLM-Based Agents as Customer Digital Twins for Conjoint Analysis', arXiv preprint arXiv:2604.22756. Available at: https://doi.org/10.48550/arXiv.2604.22756 (Accessed: 18 August 2026).

Ye, Z. and Yoganarasimhan, H. (2026) 'Rectification Difficulty and Optimal Sample Allocation in LLM-Augmented Surveys', arXiv preprint arXiv:2604.17267. Available at: https://doi.org/10.48550/arXiv.2604.17267 (Accessed: 18 August 2026).

Zhang, Y. et al. (2025) 'ConjointNet: Enhancing Conjoint Analysis for Preference Prediction with Representation Learning', arXiv preprint arXiv:2503.11710. Available at: https://doi.org/10.48550/arXiv.2503.11710 (Accessed: 18 August 2026).

Explore the idea

Let’s talk

Invisible forces shape your world — until you hire Latenta®

Contact