In-the-Moment Customer Experience Research: Catch the Answer Before Memory Rewrites It
Article M3-11
Asking people in the moment always beat asking them to remember. It lost in practice, because the pings annoyed people and someone had to hand-code every photo and voice note. AI has now cleared one of those two costs, not both.
In brief
People are poor narrators of their own past: ask a week later how an ad felt or how often they snacked, and the answer is affected by mood, story, and forgetting. Momentary methods avoid that problem by pinging a person during real life and asking right then. In theory this always outperformed recall; in practice it rarely got used, for two reasons, the fixed ping that annoyed people until they stopped responding, and the manual coding of every photo, voice note and video. Since 2022, one of those costs has dropped sharply: multimodal language models can now read a momentary capture at scale, demonstrated in peer-reviewed work and shipping across several insight platforms, though no vendor accuracy claim has been independently tested. The other cost, an AI that writes the prompt itself from live context, is still a promise, alive in preprints and an adjacent field but in no shipping product and backed by no trial. The honest state is a half-solved method: the reading bottleneck is improving, the prompting bottleneck is not.
What the theory says
The theory
The method exists to fix a known problem: when you ask someone to remember an experience, you do not get the experience back. You get a reconstruction, edited by their current mood, by the story they tell about themselves, and by plain forgetting. Studies that put the two side by side keep finding only partial agreement between what people report in the moment and what they report afterward (Freeman et al., 2023; Leertouwer, Schuurman and Vermunt, 2022). The distortion is real but not a simple uniform exaggeration: recall does not reliably inflate or deflate the past in one direction, so you cannot correct for it with a fixed adjustment (Anvari et al., 2024). It is a memory-related version of the say-do gap in [M0-07]: what people report is not what they did.
Momentary methods ask the question at the moment it happens. Two names cover most of the field. Experience sampling, or ESM, prompts a person several times a day and asks what they are doing, feeling or thinking right now. Ecological momentary assessment, or EMA, is the same idea in health and clinical settings, catching a symptom or a craving as it happens rather than at the next appointment. A close cousin is mobile ethnography, the digital diary, where participants film, photograph and narrate their own lives on their phones as they go (Palmberger, 2025). All share one assumption: an answer given in the moment is better than one reconstructed later.
That assumption was supported in the literature but still failed in practice, because momentary methods carried two costs recall did not. The first was the prompt: a fixed schedule that alerts someone at set times is intrusive, and people answer less as the days go on. The second was the capture: photos, voice notes and video are rich, but for decades a trained human had to read and code every one, work that did not scale. So the method best at catching a true moment was also the slowest and most expensive to run, and most studies used recall instead.
The post-2022 change affects exactly one of those costs. Multimodal language models, which read images and audio as well as text, can now interpret a momentary capture at scale. The strongest single piece of evidence is a peer-reviewed study that used a multimodal model to analyse more than 34,000 phone screenshots, reconstruct what people were actually doing, and check it against experience-sampling reports (Guo et al., 2025); related work runs the same capture-and-read loop on a smartwatch (Le et al., 2025). This is the same machine-reading of open human material that [M3-07] covers for text, now applied to photos, voice and video. The hand-coding cost, the one that kept the method small, is the one AI is removing.
The second cost is different, and confusing the two is where most of the marketing confusion comes from. An adaptive prompt, a message a model writes on the spot from your context, is the more exciting promise: no fixed, annoying schedule, just a question tailored to the moment. That half is not yet delivered.
Controversies
The current disagreement is whether an AI that authors the prompt improves momentary measurement, or whether the phrase is just a new name for something older. Consider the newer claim first. The clearest system implements a demanding introspective method, descriptive experience sampling, by putting a language model in the interviewer's seat so it no longer needs a scarce trained expert (Carmon et al., 2026). It is a serious effort, co-developed with the method's creator, but it is a preprint that presents the system, not a result: validation is pending. The nearest supporting evidence comes from an adjacent task, where a model writes a tailored message to nudge physical activity, and adaptive wording there does appear to work (Song et al., 2025), but nudging someone to exercise is intervention, not measurement, and the two have different bars. Peer-reviewed work has now put a model inside the sampling loop for a narrower job: in a small pilot it read a person's sensed context and fed back a short memory cue, which helped them report the moment (Lu et al., 2026), the first peer-reviewed step, and it is far from the adaptive prompt the idea assumes. So the adaptive prompt is plausible and pushed from several directions at once. It is not settled.
There is a more specific question underneath: is the context-aware prompt even a post-2022 feature? The context-awareness that ships today is triggered by sensors and rules, fire a prompt when the phone detects the person has stopped moving, or entered a location, or based on earlier answers (Avicenna Research et al., 2026). That is useful and it long predates large language models. What no platform advertises is a prompt whose wording a model writes live from context. The novelty everyone is excited about is narrow, and the products that would carry it do not yet do it.
A second controversy underlies the whole method: does the prompt change the moment it is trying to measure? This is reactivity, and it is not just a theoretical concern, a 2026 study found that repeatedly asking people about their drinking changed the behaviour (Kalina et al., 2026). The standing review treats reactivity as real but manageable rather than disqualifying (Stone, Schneider and Smyth, 2023). The unresolved worry for the AI version is the direction of the effect: a prompt that is more engaging and more tailored could plausibly increase reactivity rather than reduce it, and nobody has measured whether it does.
Limitations
The first limitation is the one vendors rarely mention: there is no independent test of any AI capture-reading accuracy claim. The platforms report high-quality transcription, translation and theming (dscout, Recollective and Voxpopme, 2026), but every accuracy figure comes from the vendor or a customer, with no third-party benchmark and no published head-to-head against human coders. A vendor document shows what is claimed, not that the claim is true. The peer-reviewed work shows the capability is real (Guo et al., 2025); it does not license the specific accuracy numbers a platform prints about its own product.
The second is who actually answers. Momentary methods get accurate data at each moment but reduce compliance. Response rates fall over a study's run (Howard and Lamb, 2023), and those who keep answering are not a random slice of those who started; burden and dropout are selective (Tate et al., 2024), and the features of a prompt predict whether it gets answered (Murray et al., 2023). The recall advantage can be greatly reduced by selection: a dataset accurate about each recorded moment can still be heavily biased toward people who comply, and reading the captures with AI does nothing to fix who supplied them.
A third is a design trade-off that no one has solved. Microinteraction EMA reduces the prompt to a single tap so people keep responding for months, and engagement stays high, but a one-tap answer contains much less information than a filmed diary entry (Ponnada et al., 2025). AI reading the capture does not fix that, because there is less capture to read.
Finally, the two halves are not supported by the same evidence. The reading half is peer-reviewed (Guo et al., 2025); the prompting half rests on a preprint and an adjacent-task result, not a validated one, and it should be interpreted that way.
Open questions
The most important test has not been run: no published, controlled study shows that an AI-authored adaptive prompt beats a fixed ping on either compliance or data quality. This is the basic assumption behind adaptive prompts, and it is currently unevidenced, which is not the same as false. It is untested.
Two smaller unknowns remain. Whether an AI-written prompt raises reactivity, by being more engaging, is plausible and unmeasured. And the standards have not been updated: the 2025 edition of the ICC/ESOMAR International Code adds articles on the responsible use of AI, synthetic data and transparency and keeps human oversight central (ICC/ESOMAR, 2025), but sets no rule specific to momentary methods or to an AI that writes the prompt.
So what
The practical shape is simple. Measuring in the moment beats measuring from memory, and it always did; what stopped it being the default was two costs, the annoying fixed ping and the hand-coding of every capture. AI has knocked down the second decisively and left the first almost untouched. So the useful thing you can buy today is not a smarter prompt but a machine that reads photos, voice and video at a scale a human team never could, which finally makes rich momentary data affordable to analyse.
One ethical line runs through every use below and belongs at the front. This method reaches into people's real lives and can disturb the moment it measures: a prompt engaging enough to keep people answering is also engaging enough to change what they do, and a capture of someone's actual day is more intimate than a survey response, which is why responsible frameworks for AI-augmented digital research put ethics and human oversight at the centre (Cheah, 2025). The rigorous version and the responsible version are the same: you are accountable for a number that has to survive contact with reality, and a measurement tuned to be pleasant is worth nothing as intelligence. The test to apply everywhere is whether a choice serves the decision you are making, or the answer someone already wanted.
For research practice
Use momentary capture where recall is the enemy, meaning any question about feeling, effort, timing or frequency that a person cannot faithfully reconstruct later. The new affordability is real: AI reading the captures removes the bottleneck that kept this method off most projects. Treat the model's output as a first-pass coding, not a verdict, because no independent benchmark yet tells you its error rate against a human coder. Spot-check a sample by hand, always. Watch compliance harder than accuracy, since the biggest threat to a momentary dataset is not a misread photo but the fact that your most burdened participants quietly stopped answering, and AI does nothing about that. And keep the two halves of the promise straight. Paying for a model that reads captures buys you something that works. Paying for a model that writes the ping from live context buys you something that no vendor has shown works, because none of them ships it.
For companies
You will be told that AI-analysed mobile ethnography and video diaries at scale are a direct alternative to recall-based surveys. The benefit is real. Reading momentary photos, voice and video used to be expensive, but now it is cheap, so you can run richer in-the-moment studies across more of the business than before. Two warnings. First, when a platform reports how accurate its AI analysis is, ask what that number was measured against, because the published validation of the underlying capability is not the same as a vendor's claim about its own product, and no independent test of these accuracy figures was found. Second, the conversational follow-up and probing features these platforms are adding are the same craft covered in [M1-01] and [M1-02]; here they are simply reaching momentary contexts, and the questions to ask about them live there. The mistake to avoid is buying momentary data and forgetting who supplied it. A concept diary answered only by your most engaged customers does not represent your market.
For political parties
Question and prompt wording move numbers, and momentary capture is appealing because getting a voter's reaction during an ad or a debate is better than asking them the next day what they felt. The ethical line comes first here, because this is where measurement becomes manipulation. A prompt that is adaptive and engaging can steer the very reaction it claims to record, and reactivity means the act of measuring repeatedly can change behaviour (Kalina et al., 2026). A momentary study designed to produce a warm reaction will produce one, and it will tell you nothing you can act on. The rigorous move is the honest one: use momentary capture to observe a genuine reaction you have not shaped, apply the decision test to every prompt you write, and remember that a measurement designed to make a message look good will be exposed on election day. The recall problem is real and worth solving. Solving it with an instrument that plants the answer is not smart, it is a show you paid for.
For government and policy
Official statistics and public-health surveys are most affected by getting a moment right, and momentary methods are already proven to run in demanding field settings, from grief research to managing a chronic lung condition at home (Mintz et al., 2024; Miller et al., 2024). AI reading the captures lowers the cost of doing this at population scale, which is a genuine improvement. Two constraints deserve to be written into any use. Compliance is selective, and the people least likely to keep answering a stream of prompts are often exactly the ones an official survey exists to reach, so a momentary dataset can under-represent the underserved without being noticed even when each recorded moment is accurate. And momentary capture of someone's actual day is sensitive data with important consent issues, more so when a model is reading it. The appropriate approach is AI as a first screen that widens coverage, human validation on the questions and populations where a wrong reading has consequences, and a clear internal record of what was captured, how it was read, and who dropped out.
How to use this
Use the parts that work and be doubtful of the parts that do not. Use momentary capture wherever a later memory would be unreliable, and let AI read the photos, voice and video, because that is the step it genuinely makes affordable. Then apply five checks before you trust the output. Was the AI coding spot-checked against a human on a real sample, or taken on faith? When a vendor quotes an accuracy figure, what was it measured against, and by whom? Who actually kept answering, and do they look like the group you care about? Could the prompt itself have changed the moment you were trying to measure? And if a tool claims to write the ping adaptively from live context, ask to see the evidence it beats a fixed prompt, then expect not to get any, because that trial has not been published. The technology has made rich momentary measurement cheap to analyse for the first time. It has not yet made the fixed prompt unnecessary, and assuming it has is where the method starts misleading the people who rely on it.
Case studies
Indeemo (2026). A mobile-ethnography platform that does the exact move at the centre of this article. Participants capture photo, video and voice diaries as they live, and the platform uses multimodal AI to transcribe, translate across more than thirty languages, and theme the material with a timestamp trail back to the source, instead of a human coding every clip. It positions the method explicitly against recall-based research. The boundary is the honest part: it markets AI that reads the capture, not AI that writes the prompt (Indeemo, 2026). It is the reading bottleneck clearing, in a product you can point at.
A 34,000-screenshot study (2025). The strongest single proof that the reading bottleneck is real and closing comes from peer-reviewed work, not a vendor. Researchers used a multimodal language model to analyse more than 34,000 phone screenshots, reconstruct what people were actually doing on their devices, and triangulate that against experience-sampling reports (Guo et al., 2025). It demonstrates the capability the whole spine rests on, in a venue that reviewed it.
The product that does not exist (2026). The most useful negative case is an absence. No vendor ships a system that both schedules a momentary ping and has a language model author it live from context. The platforms that read captures do not schedule adaptive prompts, and the engines that schedule prompts trigger them by sensor and rule, not by a model writing the words (Avicenna Research et al., 2026; Indeemo, 2026). The adaptive-prompt half of the promise has no product behind it yet, and that gap is the honest state of the field.
References
Anvari, F., Moeck, E.K., Franco, V.R., Elson, M. and Schneider, I.K. (2024) 'The "memory-experience gap" for affect does not reflect a general memory bias to overestimate past affect', Emotion, 24(8), pp. 1950–1961. Available at: https://doi.org/10.1037/emo0001404 (Accessed: 18 August 2026).
Avicenna Research, ExpiWell, LifeData, movisensXS, SEMA3 and ilumivu mEMA (2026) Academic and clinical experience-sampling / ecological-momentary-assessment engines: fixed, random, sensor- and context-triggered momentary prompting. Available at: https://avicennaresearch.com, https://expiwell.com, https://lifedatacorp.com, https://www.movisens.com/en/products/movisensxs, https://sema3.com and https://ilumivu.com/apps/ecological-momentary-assessment-app (Accessed: 19 August 2026).
Carmon, J., Bersch, C., Fernyhough, C., Hurlburt, R.T. and Kühn, S. (2026) 'Capturing Inner Experience At Scale: An AI Interviewer Co-Developed with the Founder of a Landmark Phenomenological Method', arXiv:2607.20310 [preprint]. Available at: https://arxiv.org/abs/2607.20310 (Accessed: 18 August 2026).
Cheah, C.W. (2025) 'AI-Augmented Netnography: Ethical and Methodological Frameworks for Responsible Digital Research', International Journal of Qualitative Methods, 24. Available at: https://doi.org/10.1177/16094069251338910 (Accessed: 18 August 2026).
dscout, Recollective and Voxpopme (2026) Commercial mobile-ethnography and video-diary insight platforms: AI transcription, translation, theming, sentiment, summarisation and AI-moderator dynamic follow-ups. Available at: https://dscout.com, https://recollective.com/features and https://voxpopme.com (Accessed: 19 August 2026).
Freeman, L.K., Haney, A.M., Griffin, S.A., Fleming, M.N., Vebares, T.J., Motschman, C.A. and Trull, T.J. (2023) 'Agreement between momentary and retrospective reports of cannabis use and alcohol use: comparison of ecological momentary assessment and timeline followback indices', Psychology of Addictive Behaviors, 37(4), pp. 606–615. Available at: https://doi.org/10.1037/adb0000897 (Accessed: 18 August 2026).
Guo, L., Fu, Y., Lin, X., Xu, X. ('Orson'), Chang, Y.-J. and Hiniker, A. (2025) 'What Social Media Use Do People Regret? An Analysis of 34K Smartphone Screenshots with Multimodal LLM', Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, pp. 1–23. Available at: https://doi.org/10.1145/3706598.3713724 (Accessed: 18 August 2026).
Howard, A.L. and Lamb, M. (2023) 'Compliance trends in a 14-week ecological momentary assessment study of undergraduate alcohol drinkers', Assessment, 31(2), pp. 277–290. Available at: https://doi.org/10.1177/10731911231159937 (Accessed: 18 August 2026).
ICC/ESOMAR (2025) ICC/ESOMAR International Code on Market, Opinion and Social Research and Data Analytics. 5th edn. Available at: https://iccwbo.org/news-publications/business-solutions/iccesomar-international-code-market-opinion-social-research-data-analytics/ (Accessed: 19 August 2026).
Indeemo (2026) Mobile ethnography: multimodal AI transcription and translation (30+ languages), timestamp-traceable thematic analysis, highlight reels and packaging OCR of participant-captured photo/video/voice momentary diaries. Available at: https://www.indeemo.com/mobile-ethnography (Accessed: 19 August 2026).
Kalina, E., Russell, M.A., Rodríguez, G.C., Leeman, R.F. and Scaglione, N.M. (2026) 'Behavioral reactivity to ecological momentary assessment of alcohol and sexual assault protective behavioral strategies', Experimental and Clinical Psychopharmacology, 34(3), pp. 233–242. Available at: https://doi.org/10.1037/pha0000807 (Accessed: 18 August 2026).
Le, H., Potter, V., Lakshminarayanan, R., Mishra, V. and Intille, S. (2025) 'Feasibility and utility of multimodal micro ecological momentary assessment on a smartwatch', Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, pp. 1–22. Available at: https://doi.org/10.1145/3706598.3714086 (Accessed: 18 August 2026).
Leertouwer, I., Schuurman, N.K. and Vermunt, J.K. (2022) 'Are retrospective assessments means of people's experiences? Accounting for interpersonal and intrapersonal variability when comparing retrospective assessment data to ecological momentary assessment data', Journal for Person-Oriented Research, 8(2), pp. 52–70. Available at: https://doi.org/10.17505/jpor.2022.24855 (Accessed: 18 August 2026).
Lu, T., Liang, R.-H., Hu, J., Van Gorp, P. and Markopoulos, P. (2026) 'ContextESM: Using LLM-Generated Summaries from Contextual Data to Support Experience Sampling Methods', in Pervasive Computing Technologies for Healthcare (EAI PervasiveHealth 2025), Lecture Notes of the Institute for Computer Sciences, Social Informatics and Telecommunications Engineering (LNICST). Cham: Springer Nature Switzerland, pp. 53–67. Available at: https://doi.org/10.1007/978-3-032-27582-0_4 (Accessed: 19 August 2026).
Miller, S., Nichols, M., Teufel, R., Silverman, E. and Walentynowicz, M. (2024) 'Use of ecological momentary assessment to measure dyspnea in COPD', International Journal of Chronic Obstructive Pulmonary Disease, 19, pp. 841–849. Available at: https://doi.org/10.2147/copd.s447660 (Accessed: 18 August 2026).
Mintz, E.H., Toner, E.R., Skolnik, A.M., Pan, A., Frumkin, M.R., Baker, A.W., Simon, N.M. and Robinaugh, D.J. (2024) 'Ecological momentary assessment in prolonged grief research: feasibility, acceptability, and measurement reactivity', Death Studies, 50(3), pp. 482–494. Available at: https://doi.org/10.1080/07481187.2024.2433109 (Accessed: 18 August 2026).
Murray, A.L., Brown, R., Zhu, X., Speyer, L.G., Yang, Y., Xiao, Z., Ribeaud, D. and Eisner, M. (2023) 'Prompt-level predictors of compliance in an ecological momentary assessment study of young adults' mental health', Journal of Affective Disorders, 322, pp. 125–131. Available at: https://doi.org/10.1016/j.jad.2022.11.014 (Accessed: 18 August 2026).
Palmberger, M. (2025) 'The digital diary: a mobile, multimodal, and participatory method and part of digital ethnography', International Journal of Qualitative Methods, 24. Available at: https://doi.org/10.1177/16094069251329262 (Accessed: 18 August 2026).
Ponnada, A., Wang, S.D., Li, J., Wang, W.L., Dunton, G.F., Hedeker, D. and Intille, S.S. (2025) 'Longitudinal user engagement with microinteraction ecological momentary assessment (μEMA)', Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, 9(3), pp. 1–27. Available at: https://doi.org/10.1145/3749541 (Accessed: 18 August 2026).
Song, H., Hofer, D., Islambouli, R., Hawkins, L., Bhattacharjee, A., Hassanzadeh, Z., Smeddinck, J., Franklin, M. and Williams, J.J. (2025) 'Tailored Behavior-Change Messaging for Physical Activity: Integrating Contextual Bandits and Large Language Models', arXiv:2506.07275 [preprint]. Available at: https://arxiv.org/abs/2506.07275 (Accessed: 18 August 2026).
Stone, A.A., Schneider, S. and Smyth, J.M. (2023) 'Evaluation of pressing issues in ecological momentary assessment', Annual Review of Clinical Psychology, 19(1), pp. 107–131. Available at: https://doi.org/10.1146/annurev-clinpsy-080921-083128 (Accessed: 18 August 2026).
Tate, A.D., Fertig, A.R., de Brito, J.N., Ellis, É.M., Carr, C.P., Trofholz, A. and Berge, J.M. (2024) 'Momentary factors and study characteristics associated with participant burden and protocol adherence: ecological momentary assessment', JMIR Formative Research, 8, p. e49512. Available at: https://doi.org/10.2196/49512 (Accessed: 18 August 2026).
Explore the idea
Let’s talk
Invisible forces shape your world — until you hire Latenta®
Contact