Here is one survey question. We sent it into three languages by machine, and every translation reads perfectly. The trouble is what the question now measures. Bring each one back to English to find out.
Wie sehr vertrauen Sie diesem Verkaeufer?
この販売者をどの程度信用していますか
ما مدى أمانة هذا البائع
All three read fluently. Two quietly changed what the question measures. Creditworthiness and honesty are close to trust, but they are not trust, and a respondent answers the question in front of them, not the one you wrote.
Illustrative trust scores: German 8.4 · Japanese 8.1 · Arabic 7.9. A clean ranking, built on three different questions.
Fluency is not equivalence. Before you compare an average score across these markets, an invariance test has to show the question still measures the same thing in each language. Machine translation gives you the fluent draft, not that proof.
Illustrative example, not recorded model output. The translations are simplified to make the drift easy to see, and a Webflow embed cannot make a live model call. Recorded 2026-08-18.