Multilingual Research and Translation: The Machine Speaks Fifty Languages. Your Question Might Not.
Article M1-07
Machine translation has made multilingual research fast and cheap, and one 2024 benchmark found it now handles culture-bound wording better than the neural tools it replaced. What it cannot promise is that a question still measures the same thing once it crosses into another language.
In brief
Machine translation is now cheap, wide, and more accurate than it has ever been, because large language models read context and meaning better than older tools did. No one has shown that the machine carries the deliberate meaning a good researcher builds into a questionnaire, so the check is still yours. A skilled researcher chooses each word and cultural reference on purpose, so the question measures what it is meant to measure. Fluent output does not mean meaning survived the crossing.
The honest move is a split. Let the machine do the fast full draft, including answer options and skip logic. Then bring a native speaker in to review the result rather than retranslate it. The machine translates well, but no one has shown it carries the built-in meaning. The check is still yours.
What the theory says
The theory
A good questionnaire carries more than its words. A skilled researcher chooses each word and cultural reference on purpose. The question measures the exact thing it is meant to measure. A question measures something called a construct: trust, or satisfaction, or brand loyalty. When a French respondent and a Japanese respondent both pick agree on the same scale, the two answers are only worth comparing if the question meant the same thing to each of them. That deliberate meaning is the research value. It is exactly what a translation has to carry across.
Getting a question to mean the same thing in a new language was slow, expensive, human work. The standard method is called TRAPD. It includes back-translating the result to catch where meaning had drifted (Walde and Völlm, 2023). That cost is why most projects that needed many languages could only afford a few (Benlidayi, 2024).
Here is what changed after 2022. Translation now happens inside the instrument, in real time, by machine. General-purpose machine translation is built straight into survey platforms as a reviewable draft. Qualtrics ships an integration with Google's translation service that covers dozens of language codes (Qualtrics, 2026). A newer class of AI interviewer goes further: it runs the whole conversation in the respondent's language, with no bilingual moderator and no separate translated questionnaire at all (User Intuition, 2026). The translation bottleneck that shaped forty years of practice suddenly has a machine-shaped way around it. But a machine translation can read perfectly in the target language and still measure a different construct, because fluent output is exactly what makes a shifted question look correct. Whether the machine removes the bottleneck and also carries the meaning the bottleneck was there to protect is the question.
There has always been a way to check whether translated versions are truly comparable. It is called measurement invariance testing. In plain terms: does the same score mean the same thing in every language, so that an average can be compared across countries and mean something.
Controversies
The live controversy is about the machine's capability, and it is more encouraging than you would expect. The assumption is that a machine must be cruder than a human translator on anything cultural. One 2024 benchmark cuts against that. Yao et al. (2024) tested machine translation on culture-specific content. Large language models handled it better than older neural translators, especially for items with no clean equivalent in the target culture. There is a catch. The standard automated scores used to grade translation quality did not show the improvement. Those scores measure word-level and grammatical adequacy, not whether the cultural sense survived. The gain is real and, by the usual instruments, invisible. This is one benchmark, not a settled capability, but it points the right way. The machine is better than its predecessor at the hardest part of the job. The usual quality checks are looking in the wrong place to see it.
The referee itself is under revision too. Two peer-reviewed 2025 papers argue that strict measurement invariance is rarely achievable or the right test for comparative research (Fischer et al., 2025; Kusano, Napier and Jost, 2025). The yardstick was already contested before any machine touched it, though in a literature this young the rebuttals have not been published yet.
Limitations
The most important limitation is an absence. Every vendor advertises native multilingual moderation, and none of them publishes a validation showing the results are actually comparable across languages. User Intuition, for instance, supports its same depth in every language claim with an internal count of interviews conducted, not with a comparability test (User Intuition, 2026). An interview count answers how much, not whether the results are comparable. Those are different questions. This is a category-wide gap, not one company's oversight. The multilingual-by-default claim sits on vendor self-report. No independent measurement validation of these platforms has been published.
There are also ways a machine translation can fail that have no human version. Goldfarb-Tarrant, Ross and Lopez (2023) showed that moving a model across languages can worsen bias rather than preserve behaviour. Cross-lingual transfer is not the neutral copy it looks like. A 2026 preprint on adapting a model to Hindi found that cultural adaptation pushed the model toward sycophancy. It began telling respondents what they seemed to want to hear. The pull shifted from one language to the next (Sattigeri, 2026). That is a single-author, uncited preprint. It illustrates a failure mode rather than showing how common it is. Still, it is a failure mode no human back-translation was ever built to catch.
Two more limitations are about the rules the machine is quietly stepping past. The ISO 20252 standard for market, opinion and social research requires translation by people with mother-tongue competence (International Organization for Standardization, 2019). That means humans. Nothing yet certifies a machine to do that job for a research instrument. The no source language needed pitch steps past a clause still on the books. The audit trail the machine is meant to replace was already thin. Sadeghzadeh and Jadesi (2025) found real gaps in how survey translations are documented. The human process the machine is measured against was itself patchy about showing its work. That does not excuse an unlogged machine translation. It means the honest comparison is between two imperfect records, not between a machine and a flawless paper trail.
Open questions
The largest open question is the one the whole subject turns on, and it has no answer. No study has tested, head to head, whether in-instrument machine translation preserves or breaks comparability across many language pairs. The test would measure against the classic TRAPD workflow on the same instrument. A targeted search returned nothing on point. The closest evidence is nearby. Haavisto and Welsch (2024) show a model can reach useful quality when it evaluates its own translation step. Kunst and Bierwiaczonek (2023) offer early best practices for using AI translations in cross-cultural work. Albuquerque et al. (2026) compare computational and human translation directly on an English to Brazilian Portuguese instrument. All three sit at the adaptation stage rather than at a fielded comparability test. None has been replicated. Each is a single study of a single workflow.
Two smaller gaps follow. The first is equity by language resource level. There are benchmarks that score raw translation adequacy for low-resource languages. The largest, FLORES-200, covers 200 of them (NLLB Team, 2022). Adequacy is not comparability. Whether an AI interview holds its depth as you move from a well-resourced language to a poorly-resourced one, on a real instrument, is unmeasured. The second gap is whether machine failure modes even show up when comparability is tested. The Yao benchmark hinted that the machine's gains and losses sit on an axis the usual metrics ignore. A shift the machine introduces might slip past a standard comparability check unseen. Nobody has mapped machine failure onto that testing process. Until that head-to-head test exists, the responsible position is plain. In-instrument machine translation is a fast, capable drafting tool. Its comparability across languages has not been demonstrated.
So what
The machine drafts, and you keep the review. That is the whole practical answer. The machine has genuinely removed the cost that kept multilingual research small. On the hardest part, culture-bound wording, one benchmark says it can beat the tools it replaced. What it has not done is show that a question survives the crossing intact. The useful stance is neither to refuse the machine nor to trust its multilingual output as finished. Let the machine produce the fast full draft, including answer options and skip logic. Then spend the time it saved on a review you own. A native speaker checks that the instrument still makes sense in their language. They are reviewing a finished draft, not translating from scratch. The job is shorter and simpler.
For research practice
For a researcher the difference is where to spend the human hours the machine just freed up. Use in-flight translation to draft the entire instrument fast: questions, answer options and logic. Then put a native speaker on the items that carry the risk. The culturally loaded wording and the constructs with no clean equivalent in the target language get reviewed, not retranslated. Before you report any average score across countries, check that a difference between countries is a real difference and not the translation talking. The mechanics of the AI interview itself, and how the adaptive probe works, are covered in the AI moderated interview. This article is about whether its output means the same thing in every language.
For companies
For a buyer the pitch is one study in fifty markets at the price of one, and the reach is real. The catch is specific to cross-language work. Global brand and concept studies live on the assumption most at risk here. Trust, premium or ease of use mean the same thing to a shopper in Sao Paulo and one in Seoul. An ambiguous cross-language question does not fail loudly. It returns a clean-looking country ranking built on respondents who answered subtly different questions. The error never shows up in the deck. When a vendor says its AI-moderated respondents deliver the same depth in every language, ask what was actually measured to support it. The reassuring part is that the rigorous move and the commercial move are the same. A cross-market comparison you have actually validated is worth acting on. One you merely assumed gets found out in the market where the translation quietly drifted.
For political parties
For a party the difference is accountability, because this is where the temptation is sharpest. Multilingual polling makes it cheap to survey language communities a campaign could never afford to reach separately. That is a real gain for understanding a coalition. It also makes it cheap to produce numbers that look comparable across those communities and are not. A party that reports uniform support across its language groups, on the strength of machine-translated instruments nobody reviewed, has manufactured a consensus rather than measured one. The rigorous path is to use the machine to reach more communities. Then check that a loaded question, about fairness, say, or national pride, is not landing as a stronger or weaker statement in one language than another. Those are exactly the culture-bound items the machine is least safe on. A number you cannot defend as comparable is not intelligence about your coalition. It is a story you told yourself, and it will mislead the campaign that trusts it.
For government and policy
For government the stakes are highest, because these numbers set budgets and target services. The populations least well served by a drifting translation are often the ones the survey exists to reach. Presenting non-equivalent measures as if they were equivalent is not a technical slip here. It produces an official figure that guides a real decision about real people. The tools are ahead of the rulebook. The audit trail matters more here than anywhere. An agency leaning on in-flight translation should keep a clear internal record of which instruments were machine-translated and which were validated. Nobody downstream should mistake a fast draft for a checked one. Where this shades into disclosure duties and the emerging rules on AI in research, that ground belongs to the research ethics code and the AI Act for research.
How to use this
Treat fluent machine translation as a fast draft, not a finished, comparable instrument. Before you compare any answer across languages, ask three questions. First, was the translation left to the machine alone, or did a person with real command of the target language and culture review the items that carry loaded or culture-bound meaning? Second, are you about to compare average scores across languages, and if so, has that comparison actually been checked rather than assumed? Third, is your record clear enough that someone could later see which numbers were validated and which were merely trusted? Two neighbouring questions sit outside this article. Whether respondents should be told a machine is interviewing them is covered in machine-interviewer candour. The deeper trap that a model gives you a competent reading rather than a real one is the same problem set out in what the machine actually knows. This article owns the cross-language question. The technology has finally made multilingual research affordable at scale. That is a real advance over a world where most projects could reach only a handful of languages. What it has not done is make the comparability check optional.
Case studies
User Intuition (2026). This is the sharpest multilingual by default claim in the market. The platform advertises native AI moderation in more than 50 languages, several levels of conversational follow-up, and the same interviewing depth in every language, with no bilingual moderators in the loop (User Intuition, 2026). The depth claim is backed by an internal tally of interviews conducted on the company's own page. It is the clearest example of the pattern here: a genuine capability paired with a confident equivalence claim and no published test behind it.
Qualtrics and Google Translate (2026). The commercial form of translation lives inside the instrument is a survey platform with machine translation built directly into the fielding tool (Qualtrics, 2026). What makes it a useful case is that the platform's own documentation undercuts the easy reading of its own feature. It warns that machine translation is error-prone and should not be the final version shown to participants without review. The tool that made in-survey machine translation routine also told its users, in the manual, not to trust it unreviewed. That warning is the whole argument here, stated by the vendor.
Nielsen Norman Group (2026). The closest thing to an independent evaluation is a hands-on study of AI-moderated interviewing by the Nielsen Norman Group. It put two AI interviewers in front of ten research professionals and assessed how they performed (Rosala, 2026). It is a valuable read. It measures a different axis than the one here. The evaluation is about usability and interview quality, not about whether a construct stays equivalent when the same interview runs in another language. That is not a flaw in the study. It is a sign of how new the specific question is. The most credible outside evaluation available still does not measure cross-language comparability. Nobody has built the test for it yet.
References
Albuquerque, M.R. et al. (2026) 'Building Bridges Between Computational Methods and Human Translation: An English to Brazilian Portuguese Application', Journal of Cross-Cultural Psychology, 57(5), pp. 829–845. Available at: https://doi.org/10.1177/00220221261418605 (Accessed: 18 August 2026).
Benlidayi, I.C. (2024) 'Translation and Cross-Cultural Adaptation: A Critical Step in Multi-National Survey Studies', Journal of Korean Medical Science, 39(49), e336. Available at: https://doi.org/10.3346/jkms.2024.39.e336 (Accessed: 18 August 2026).
Fischer, R. et al. (2025) 'Why We Need to Rethink Measurement Invariance: The Role of Measurement Invariance for Cross-Cultural Research', Cross-Cultural Research, 59(2), pp. 147–179. Available at: https://doi.org/10.1177/10693971241312459 (Accessed: 18 August 2026).
Goldfarb-Tarrant, S., Ross, B. and Lopez, A. (2023) 'Cross-lingual Transfer Can Worsen Bias in Sentiment Analysis', Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 5691–5704. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.346 (Accessed: 18 August 2026).
Haavisto, O. and Welsch, R. (2024) Questionnaires for Everyone: Streamlining Cross-Cultural Questionnaire Adaptation with GPT-Based Translation Quality Evaluation. arXiv:2407.20608. Available at: https://arxiv.org/abs/2407.20608 (Accessed: 18 August 2026).
International Organization for Standardization (2019) ISO 20252:2019 Market, opinion and social research, including insights and data analytics: Vocabulary and service requirements. Geneva: ISO. Available at: https://www.iso.org/standard/73003.html (Accessed: 18 August 2026).
Kunst, J.R. and Bierwiaczonek, K. (2023) Utilizing AI questionnaire translations in cross-cultural and intercultural research: Insights and recommendations. PsyArXiv/OSF preprint. Available at: https://doi.org/10.31234/osf.io/sxcyk (Accessed: 18 August 2026).
Kusano, K., Napier, J.L. and Jost, J.T. (2025) 'The Mismeasure of Culture: Why Measurement Invariance Is Rarely Appropriate for Comparative Research in Psychology', Personality and Social Psychology Bulletin, 52(8), pp. 2578–2595. Available at: https://doi.org/10.1177/01461672251341402 (Accessed: 18 August 2026).
NLLB Team (2022) FLORES-200: evaluation benchmark for low-resource and multilingual machine translation (No Language Left Behind). Meta AI / Facebook Research. Available at: https://github.com/facebookresearch/flores/blob/main/flores200/README.md (Accessed: 18 August 2026).
Qualtrics (2026) Translate Survey [support documentation]. Available at: https://www.qualtrics.com/support/survey-platform/survey-module/survey-tools/translate-survey/ (Accessed: 18 August 2026).
Rosala, M. (2026) AI-Moderated Interviews: If, When, and How to Use Them. Nielsen Norman Group, 30 January. Available at: https://www.nngroup.com/articles/ai-interviewers/ (Accessed: 18 August 2026).
Sadeghzadeh, M. and Jadesi, N.N. (2025) Lost in Documentation: Professional Norms and the Gaps in Survey Translation Transparency. Research Square preprint. Available at: https://doi.org/10.21203/rs.3.rs-7326457/v1 (Accessed: 18 August 2026).
Sattigeri, S. (2026) Extending Beacon to Hindi: Cultural Adaptation Drives Cross-Lingual Sycophancy. arXiv:2602.00046. Available at: https://arxiv.org/abs/2602.00046 (Accessed: 18 August 2026).
User Intuition (2026) Multilingual Research [company webpage]. Available at: https://www.userintuition.ai/platform/multilingual-research/ (Accessed: 18 August 2026).
Walde, P. and Völlm, B.A. (2023) 'The TRAPD approach as a method for questionnaire translation', Frontiers in Psychiatry, 14, 1199989. Available at: https://doi.org/10.3389/fpsyt.2023.1199989 (Accessed: 18 August 2026).
Yao, B. et al. (2024) 'Benchmarking Machine Translation with Cultural Awareness', Findings of the Association for Computational Linguistics: EMNLP 2024, pp. 13078–13096. Available at: https://doi.org/10.18653/v1/2024.findings-emnlp.765 (Accessed: 18 August 2026).
Explore the idea
Let’s talk
Invisible forces shape your world — until you hire Latenta®
Contact