Choosing a Conjoint Engine: The Choice Engine That Fits Better and Reads Worse
Article M4-09
Conjoint studies have run on the same statistical engine for twenty years. A new wave of neural models fits the same data more closely, and, run without care, will confidently tell you what customers would choose in a world none of them were ever shown. The catch is that the version you can actually buy does not exist yet.
In brief
Conjoint and choice studies show people a series of options and watch which they pick, then work backwards to how much each feature is worth. For twenty years that backward step has run on one kind of engine, a hierarchical Bayes model that hands every respondent a readable set of weights. Since 2022 a different engine has arrived from the research literature: a neural network that learns the choice pattern directly, and, in the newest work, an amortised version that is trained once and then reused for fast evaluation. The neural engine fits the data more closely and picks up feature interactions the old one has to be told about in advance. It also carries a specific weakness. Because its inner workings are hard to read, it can extrapolate with confidence into feature and price combinations no respondent ever saw, and return a share number that looks authoritative and has no traceable basis. Two things keep this honest. The evidence is almost all 2025 and 2026 preprints, and no vendor sells the neural engine yet. Today it is a real development in papers and code, not a product you can buy.
What the theory says
The theory
A choice study exists to measure how much each feature of a product moves a buyer: the brand, the size, the warranty, and above all the price. Asking people directly does not work well, because they will say they care about quality and then buy on price (Hensher, 2010). Conjoint analysis asks indirectly instead. It shows a respondent a series of made-up options, each a different bundle of features at a different price, and records which one they pick each time. From the pattern of picks you can work backwards to the weight each feature carried, without ever asking the respondent to name it.
The maths underneath is the random utility model. Each option carries a score, its utility, built by adding up the weight of its features. People tend to choose the option with the highest score, with a random wobble that accounts for the fact that the same person does not always choose the same way. The weights are what the study is really after. In the trade they are called part-worths, and a good study gives every respondent their own set, because the point of the exercise is that people differ.
Producing those individual weights is the job of the estimation engine, and here the industry has settled on one answer. Hierarchical Bayes is the standard. It estimates a separate set of weights for each person while borrowing information across the whole sample, so that a respondent who only made a handful of choices still gets a sensible, stable set of numbers pulled toward the group. The commercial platforms that dominate conjoint work run on this engine. Sawtooth Software, whose documentation describes choice-based conjoint built on exactly this hierarchical Bayes approach, is the reference tool most practitioners learn on. The output is readable. You can look at a respondent's weights and say what they valued, and you can trust the model when it fills a small gap because it fills it in a disciplined, pre-specified way.
The new work argues with this incumbent. The challenger replaces the hand-built utility formula with a neural network, a flexible function that learns the mapping from features to choices directly from the data, without being told in advance which features interact.
Neural networks applied to choice data are not themselves new. The change after 2022 is specific. Two things are genuinely new.
The first is a proof that the flexible engine need not abandon the theory. Aouad and Désir (2025), the one peer-reviewed study anchoring this whole shift, showed that a neural network can represent any model in the random utility family and still obey its rules, and they derived a bound on how wrong it can be on choices it has not seen. That bound matters because it puts a mathematical ceiling on the engine's worst habit. Their model matched or beat both classic choice models and general machine-learning methods on real datasets.
The second is the amortised turn. Huch and Keane (2026), in a working paper circulated through the National Bureau of Economic Research, take the expensive part of fitting these models and do it once. They train a neural network to stand in for the slow calculation, and then reuse that trained network to evaluate new data quickly and deterministically. Amortised means the heavy cost is paid up front and then spread across many later uses. They also show the estimates you get this way remain statistically well-behaved even when the neural stand-in is imperfect, which is the property that would let a practitioner trust it. This is the engine the article is about: fast, flexible, trained once, and reusable.
Using a language model to supply extra preference data and cut the number of real respondents is a different subject, covered in conjoint with a model in the loop. This article is about the estimation engine, the maths that turns choices into weights, not about where the choice data comes from.
Controversies
The live argument is not whether the neural engine fits better. It is what that better fit costs, and whether the cost can be paid down.
The clearest statement of the cost comes from Villarraga and Daziano (2025). A flexible network buys its closer fit with weights you cannot read and with unstable behaviour once you push it past the range of the data it was trained on. On a small survey, which is most surveys, it can overfit, and it carries no honest sense of its own uncertainty, so it does not know when it is guessing.
What makes their paper a controversy rather than a verdict is that they also offer a fix. They build a version that leans on sensible prior assumptions when the data is thin, so the engine falls back to something defensible instead of inventing. The field treats the trade as a problem to be solved, not as a reason to walk away from neural engines.
The second open dispute is whether the neural engine's advantage holds up on the thing choice studies are actually bought for. Aggregate fit is the easy target. What a pricing team needs is individual-level differences between customers and believable predictions for feature combinations that were never tested. Acharya, Hainmueller and Xu (2026) argue their hybrid approach recovers the person-to-person differences that simple averages wash out, by keeping a readable core and adding neural flexibility around it. That is a promising claim, and it is theirs alone so far. No independent study has yet run their hybrid engine and the hierarchical Bayes incumbent side by side on the same data and reported which a buyer should trust.
It helps to separate two things that get merged. Machine-learning models fit choice data more tightly than the classic logit: that is old news, established before the 2022 line this article draws. Wang et al. (2021) made the statistical case, and Wong and Farooq (2021), with a model called ResLogit, built an early hybrid of a neural network and a choice model. The frontier claim is not about fit at all. It is about whether the flexible engine can be made readable and honest about its own limits, which is a harder and newer question than raw accuracy.
One more architecture is worth naming for what it attempts. DeepHalo (Zhang et al., 2026) is a neural choice model built to keep certain context effects under deliberate control rather than letting the network do whatever fits, which is the same instinct as the hybrids: flexibility you can still steer. It is a 2026 preprint with no independent uptake yet.
Limitations
The evidence is the first limit. Of the papers describing the new engine, only Aouad and Désir (2025) is peer-reviewed. Everything else that carries the argument is a 2025 or 2026 preprint or working paper, including the decisive amortised result from Huch and Keane, which, at the time of writing, had accrued no citations because it is too new to have any. The work is promising and fast-moving, not a settled method a practitioner can lean on today.
The second limit is commercial. No platform a client can actually buy runs the neural or amortised engine. Every platform a client can actually buy still runs hierarchical Bayes. The neural engine lives in academic papers and open-source code. The trade this article describes is real in the literature, not a choice a buyer faces through a vendor. Anyone told otherwise by a sales deck should ask to see the engine named and the validation published.
The third limit concerns the numbers. Most speed and accuracy figures circulating in this area do not measure what a buyer would want measured. Vendor speed multipliers describe faster fitting of the old model, not the neural engine beating the old one on accuracy. Where a neural paper does report a gain, it is the authors reporting on their own model. ConjointNet (Zhang et al., 2025) reports preference-prediction gains of more than five percent over classic conjoint, a real and specific result, but it is self-reported, measured on two public datasets, and not independently replicated. Read every accuracy claim in this field as the model's authors describing their own model until someone disinterested checks it.
The deepest limitation is the extrapolation worry: the neural engine may return a confident share outside the data, and that share can be fiction. Asked about a price or feature combination well outside anything a respondent saw, it will not refuse. This failure mode is argued carefully by Villarraga and Daziano (2025) and by Acharya, Hainmueller and Xu (2026), and it is bounded in theory by the error ceiling in Aouad and Désir (2025). What it is not, yet, is measured head-to-head. No study runs the neural engine and hierarchical Bayes on the same study, pushes both into invented territory, and shows the neural one going wrong as its main result. The danger is well reasoned and theoretically capped. It is not a demonstrated headline finding, and the article treats it that way.
Open questions
Can a head-to-head test settle the counterfactual worry? Someone needs to take one real choice study, estimate it with both engines, ask both for shares on feature and price combinations no respondent was shown, and report which engine stayed sensible. Everyone in this literature gestures at that experiment. No one has published it. Until someone does, the case that a neural engine will mislead you on counterfactuals is inference from theory, not a result you can point to.
Does the amortised engine generalise across studies, not just across respondents within one study? Huch and Keane (2026) show the train-once-reuse mechanism works. Whether a network trained on one market's choices can be reused on another market's, which is where the cost saving would really pay off, is not established for market-research practice.
Do the readable-utility hybrids deliver on their promise at scale? Acharya, Hainmueller and Xu, along with the other RUM-compatible networks now appearing (Bagheri, Ghasri and Barlow, 2025; Feng et al., 2024), claim to give both readability and neural flexibility at once. Each rests on its authors' own validation. None has been reproduced by an outside team on production-sized data, so the promise is credible and unconfirmed.
So what
A choice study can be estimated by one of two engines. The incumbent is readable and disciplined: it fills gaps conservatively. The challenger fits the data more closely and finds interactions nobody specified, but it can invent confidently once pushed past the data. The challenger is not yet for sale. The immediate decision is therefore not which engine to run but how to read the results either engine gives you, and how to spot the moment a flexible model has stopped reporting and started imagining. These numbers set prices and justify launches, so that reading is an accountability question before it is a technical one.
For research practice
Treat the estimation engine as part of the finding, not a black box behind it. When you commission or run a choice study, know which engine produced the shares and how far the questions pushed beyond the options respondents actually saw. Predictions inside the tested range are the trustworthy part of any conjoint, on either engine. The risk is specific: the extrapolated number, the price point or bundle nobody was shown, is where a flexible neural engine can fail quietly while hierarchical Bayes stays sensible. When the neural engine becomes buyable, demand a readable account of what it learned and a stated range outside which its shares are not to be trusted. The strongest current use of this literature is defensive: it tells you which number in a conjoint deck to interrogate, and that is the extrapolated one.
For companies
Before the pull of a neural engine becomes a purchase, hold it to the standard the decision deserves. A neural engine promises to squeeze more out of the same survey and to answer questions about untested products, exactly what a pricing team wants. A share prediction that justifies a price is only as good as its origin, and a confident number from an unreadable model asked about a product no respondent evaluated is not evidence: it is a guess wearing a decimal point. A model tuned until it produces the share the business hoped for is a mirror, and it gets found out when the product ships to the real market at the real price.
The useful question to put to any vendor selling a smarter conjoint engine is not how much better it fits. It is where its predictions stop being reliable, and whether the vendor will tell you when a question has taken it past that line. A model that always returns a number, with no line at all, has given you the problem rather than the feature.
For political parties
In political choice-style experiments, the temptation is to run the specification that returns the coalition you wanted, and an unreadable engine makes that temptation cheaper and harder to catch. It can produce a favourable number for an untested bundle with no visible reasoning to challenge. Choice-style experiments now test messages and policy bundles, not just products, and the flexible engine can score combinations never fielded. Use the flexible engine to narrow a wide field of options quickly, then confirm the survivors on real respondents actually shown those options. A projected level of support for a policy nobody was asked about is not a reading of the electorate; it is a hypothesis, not a finding. Treating it as a finding is how a campaign ends up believing its own preferred answer right up until the vote contradicts it.
For government and policy
Choice experiments inform real decisions in government, from how public services are valued to how much people will pay for an environmental gain, and willingness-to-pay numbers from these studies carry legal and budgetary weight. A number that sets policy must be traceable to the responses that produced it, which an unreadable engine cannot guarantee once it starts extrapolating. The classic hierarchical Bayes engine has a real advantage here that has nothing to do with accuracy: you can show your working. For an agency, the defensible position is to treat the newer engines as what they are, a promising research literature resting largely on preprints, and to require of any future tool the same thing you would require of a human analyst: an account of how it reached the number and a clear statement of where its estimates should not be trusted. The related question of how a model should express its own uncertainty is a subject in its own right, taken up in reporting what a model does not know.
How to use this
Keep three things straight. First, this is a live research frontier, not a settled method, so the value now is in reading conjoint results more sharply, whatever engine made them. Second, when a deck shows a share for a price or bundle nobody was shown, that extrapolated number is the one to question first. Third, when these engines do reach the market, judge them on readability and honest limits, not on fit. An engine that fits beautifully and cannot tell you where it stops being reliable is more dangerous than the plainer engine it replaces, because it fails in the one place a choice study is least able to check it. The willingness-to-pay and stated-preference cautions that sit underneath all of this are covered in what people say they will pay; this article adds only what changes when a flexible engine, rather than a person, fills in the answer nobody gave.
Case studies
An industrial research team, ConjointNet (2025). A group of applied researchers built two neural architectures for conjoint analysis and reported preference-prediction gains of more than five percent over classic conjoint on two public datasets (Zhang et al., 2025). It is the clearest worked example of the neural engine beating the incumbent on fit, and also a clean illustration of the honest-status problem. The result is real, specific, and self-reported, measured by the model's own authors and not yet checked by anyone independent. Read it as a strong signal that the fit gain is achievable, not as a settled benchmark. The affiliation of the team is not stated on the record this article could verify, so it is described here only as an industrial research group.
A worked academic comparison (2025). Villarraga and Daziano (2025) built a Bayesian version of a deep choice model and validated it on real transport data, including New York City travel choices and a Swiss rail dataset, recovering readable trade-off rates with honest interval estimates rather than single point guesses. It is the best current example of the interpretability-versus-flexibility trade being worked on with real data rather than only argued. It also marks the limit of the current evidence. It shows a flexible engine can be made to report its own uncertainty on real choices. It does not settle whether that engine beats hierarchical Bayes on the counterfactual shares a buyer most wants, which is the open test at the centre of this whole area.
References
Acharya, A., Hainmueller, J. and Xu, Y. (2026) 'Learning preferences from conjoint data: a hybrid structural deep learning approach', arXiv preprint. Available at: https://doi.org/10.48550/arXiv.2604.10845 (Accessed: 18 August 2026).
Aouad, A. and Désir, A. (2025) 'Representing random utility choice models with neural networks', Management Science, 72(8), pp. 6686–6701. Available at: https://doi.org/10.1287/mnsc.2023.02189 (Accessed: 18 August 2026).
Bagheri, N., Ghasri, M. and Barlow, M. (2025) 'RUM-NN: a neural network model compatible with random utility maximisation for discrete choice', arXiv preprint. Available at: https://doi.org/10.48550/arXiv.2501.05221 (Accessed: 18 August 2026).
Feng, S., Yao, R., Hess, S., Daziano, R.A., Brathwaite, T., Walker, J.L. and Wang, S. (2024) 'Deep neural networks for choice analysis: enhancing behavioral regularity with gradient regularization', arXiv preprint. Available at: https://doi.org/10.48550/arXiv.2404.14701 (Accessed: 18 August 2026).
Hensher, D.A. (2010) 'Hypothetical bias, choice experiments and willingness to pay', Transportation Research Part B: Methodological, 44(6), pp. 735–752. Available at: https://doi.org/10.1016/j.trb.2009.12.012 (Accessed: 18 August 2026).
Huch, E.K. and Keane, M.P. (2026) 'Amortized inference for correlated discrete choice models via equivariant neural networks', arXiv preprint (also NBER Working Paper No. 35037). Available at: https://doi.org/10.48550/arXiv.2603.24705 (Accessed: 18 August 2026).
Sawtooth Software (no date) Choice-based conjoint (CBC) analysis. Available at: https://sawtoothsoftware.com/conjoint-analysis/cbc (Accessed: 18 August 2026).
Villarraga, D.F. and Daziano, R.A. (2025) 'Bayesian deep learning for discrete choice', arXiv preprint. Available at: https://doi.org/10.48550/arXiv.2505.18077 (Accessed: 18 August 2026).
Wang, S., Wang, Q., Bailey, N. and Zhao, J. (2021) 'Deep neural networks for choice analysis: a statistical learning theory perspective', Transportation Research Part B: Methodological, 148, pp. 60–81. Available at: https://doi.org/10.1016/j.trb.2021.03.011 (Accessed: 18 August 2026).
Wong, M. and Farooq, B. (2021) 'ResLogit: a residual neural network logit model for data-driven choice modelling', Transportation Research Part C: Emerging Technologies, 126, article 103050. Available at: https://doi.org/10.1016/j.trc.2021.103050 (Accessed: 18 August 2026).
Zhang, S., Wang, Z., Gao, R. and Li, S. (2026) 'DeepHalo: a neural choice model with controllable context effects', arXiv preprint. Available at: https://doi.org/10.48550/arXiv.2601.04616 (Accessed: 18 August 2026).
Zhang, Y., Chen, F., Hakimi, S., Harinen, T., Filipowicz, A., Chen, Y.-Y., Iliev, R., Arechiga, N., Murakami, K., Lyons, K., Wu, C. and Klenk, M. (2025) 'ConjointNet: enhancing conjoint analysis for preference prediction with representation learning', arXiv preprint. Available at: https://doi.org/10.48550/arXiv.2503.11710 (Accessed: 18 August 2026).
Explore the idea
Let’s talk
Invisible forces shape your world — until you hire Latenta®
Contact