ChatGPT-4o as an English–Nepali voice interpreter in the field: 58 % of utterances faithful, and 91 % of the worst errors when the patient is the one speaking
Six authors from the University of Washington, Georgetown, NEOMED and Kathmandu University School of Medical Sciences ran 30 real conversations — on walking paths and at roadside shops around Dhulikhel Hospital in Nepal — through ChatGPT-4o's voice mode, then had a bilingual reviewer rate all 485 translated utterances. 58.1 % of translations preserved meaning, but 14.2 % lost it — and 63 of those 69 failures, or 91.3 %, occurred when the Nepali-speaking participant was the one talking. The paper's most useful finding is not the overall rate: it is that the most inaccurate translations stayed fluent enough that the English-speaking listener would notice nothing.
The context
Language discordance — clinician and patient, or researcher and participant, not sharing a language — measurably degrades comprehension, satisfaction and quality of care. The professional interpreter remains the standard: their presence improves all three compared with ad hoc interpretation (a relative, a bilingual receptionist) or no language assistance at all. But using an interpreter presupposes budget, trained staff, integration into the health system and scheduling — four conditions that are precisely what rural areas and unscheduled encounters lack.
Hence the interest in digital tools. The previous generation chained three separate modules: speech recognition, machine translation, speech synthesis. Multimodal language models — able to process audio directly and produce speech through a single conversational interface — remove those seams and fit in a phone. ChatGPT has been explicitly proposed as a multilingual communication solution in low- and middle-income countries. Accessibility and reliability are, however, two different things: none of these systems was developed or validated as a substitute for an interpreter.
Nepali is what is called a low-resource language: few parallel corpora, little annotated audio, limited research. It also combines features that complicate machine translation — subject-object-verb structure, context-dependent pronouns, a system of honorifics (respectful tapai versus familiar timi), culturally grounded expressions, regional dialect variation. Earlier evaluations of Nepali machine translation had already identified lexical, syntactic, semantic and cultural losses.
Above all, most published evaluations concern written text: standardized sentences, fixed datasets, prepared discharge instructions, retrospective comparison with a reference translation. But the real use of a voice translator is spontaneous speech, with hesitations, repetitions, unfinished sentences, background noise and overlapping turns. That is the gap this paper fills.
The method
The setup. Cross-sectional observational study, conducted in July and August 2025 within roughly 5 kilometres of Dhulikhel Hospital, Kavrepalanchok District. Convenience recruitment in public spaces — walking paths, roadside shops — explicitly outside hospital grounds, and no current patients were approached. Inclusion criteria: 18 or older, primarily Nepali-speaking, able to consent. No exclusion on education level, literacy, technology familiarity or prior healthcare use. The protocol targeted 20 to 30 participants; 30 were enrolled. No power calculation: the study is descriptive and exploratory, and the authors say so.
The equipment. Paid OpenAI subscription, GPT-4o model, iPhone 13 Pro on iOS 16.6.1, Ncell mobile network. A fresh ChatGPT session for each participant. A strictly identical initialization prompt every time, instructing the model to translate only what was said, without adding or summarizing anything. A separate audio recording covered each full conversation.
The questions. Four, in English, finalized after pre-testing: age; whether the person uses Dhulikhel Hospital (and if so, what for and how they travelled there); what they do to stay healthy; their favourite festival and how they celebrate it. They are deliberately non-diagnostic and low-risk — the aim is to elicit natural speech, not to assess clinical decision-making. The team's Nepali-speaking collaborator attended recruitment but did not take part in the conversation: the model was alone in the loop.
The evaluation. A single bilingual reviewer, at ACTFL Advanced-Mid proficiency, re-listened to the recordings alongside the ChatGPT text transcripts and rated each of the 485 translated utterances on a three-point scale: 3 = core meaning fully preserved and naturally rendered; 2 = main idea preserved with one or two noticeable errors of grammar, word choice or phrasing; 1 = meaning significantly altered, unclear or lost. In parallel she built an inductive taxonomy of deviations — that is, derived from the data rather than from a pre-existing scheme — refined across transcripts, then applied a second time to the whole corpus with the final category set. A single utterance could receive several tags. Nineteen deviation types were identified, among them: distortion of meaning, overly formal phrasing, omission, addition, over-interpretation, number error, use of Hindi words, meaning loss, delay over three seconds, no translation produced, mid-sentence cutoff, overly casual register, inappropriate dialect, wrong output language.
Ethics. Protocol approved by the Institutional Review Committee of Kathmandu University School of Medical Sciences (IRC-KUSMS, approval no. 171/25, 16 July 2025). Written consent, separate authorization for audio recording, explicit notice that speech would be processed by ChatGPT on external servers.
The results
The overall distribution. Across 485 translated utterances, 282 (58.1 %) rated 3, 134 (27.6 %) rated 2, 69 (14.2 %) rated 1. In other words, slightly more than one utterance in seven undergoes a meaning change judged significant, and nearly one in three requires clarification.
The directional asymmetry is the central result. English to Nepali, mean rating 2.63 ± 0.53: of 257 utterances, 167 (65.0 %) rated 3 and only 6 (2.3 %) rated 1. Nepali to English, mean rating 2.23 ± 0.86: of 228 utterances, 115 (50.4 %) rated 3 but 63 (27.6 %) rated 1. Of the 69 failures in the corpus, 63 — 91.3 % — therefore occur when the participant is speaking. The standard deviation is systematically larger in that direction, for every question: it is not only worse, it is less predictable.
The most open-ended question is the worst translated. English to Nepali, the best rating goes to the age question (2.79 ± 0.48), then the hospital question (2.68 ± 0.47) and the festival question (2.64 ± 0.48); the worst to "what do you do to stay healthy?" (2.33 ± 0.63). Nepali to English, the ordering is the same but compressed downward: 2.33 ± 0.83 for the hospital, 2.26 for age and festival, and 1.85 ± 0.86 for health. The question that invites a free-form answer with no expected shape is where the system breaks down.
The failure modes. Distortion of meaning leads (65 instances, 17.1 % of assigned tags), ahead of overly formal phrasing (56; 14.7 %), omission (54; 14.2 %) and addition (44; 11.5 %). Then over-interpretation (23; 6.0 %), clarification requests (15; 3.9 %), grammar errors (14; 3.7 %), number errors (13; 3.4 %), use of Hindi (12; 3.2 %), delay over three seconds and no translation produced (10 each; 2.6 %), meaning loss (7; 1.8 %), response in English where Nepali was expected and mid-sentence cutoff (6 each; 1.6 %). Thirty-five tags (9.2 %) correspond to a participant switching to English themselves — a conversational choice, not a system error.
What this looks like in practice. The reported examples say more than the percentages. A participant states "I farm. Agricultural work"; the output is "I walk a little every day." An age of 67 becomes 47. Someone says they like all festivals; the system attributes to them a preference for Dashain and Tihar. Asked "what do you do to stay healthy?", the model instead produced an earlier question about transport to the hospital. Elsewhere respectful tapai becomes familiar timi, making the speaker sound disrespectful without any change of meaning. The Hindi word achha replaces Nepali thik cha. Conversely, the model correctly rendered the idiom sathi bhai as "friends" rather than literally.
Practical translation. For a twenty-turn interview on the patient's side, the distribution observed in the Nepali-to-English direction yields on average five to six utterances whose meaning is significantly altered, and roughly ten that are only partly faithful. The critical point is not the volume: it is that those altered utterances arrive in fluent, grammatical, plausible English, with no signal allowing the listener to spot them.
What is good
The protocol is constant where it matters, and the evaluation is done under real conditions. Same English-speaking facilitator, same device, same account, same model version, same initialization prompt, same questions, across all 30 encounters. That stability is what licenses attributing the observed variance to speech and to the system rather than to the apparatus. And it applies to spontaneous outdoor conversations, with noise, hesitation and code-switching — not to a written corpus. This is precisely the real usage condition the existing literature did not cover.
The taxonomy separates what AUC or BLEU aggregate. An overall rating on one side, nineteen deviation categories on the other: the design explicitly distinguishes errors that damage the naturalness of the conversation (overly formal register, wrong dialect, latency) from those that damage fidelity (distortion, addition, number error). That separation is the paper's reusable contribution. A single aggregate metric would have blended a bookish phrasing with an age wrong by twenty years, which do not carry the same consequences.
The authors dismantle their own most marketable finding. The directional asymmetry is the study's potential headline; they write that it is confounded with input type, since the English questions were standardized and repeated thirty times while the Nepali responses were spontaneous and varied in length and clarity. They add that they cannot distinguish an error arising in speech recognition from one arising in translation. The de-identified dataset of ratings and codes is published as a supplement. No competing interests declared, and no industry funding appears.
What is less good
One reviewer, no inter-rater reliability — and that is the entire measurement. The 485 ratings and 19 deviation categories come from one person. The authors flag it, but the scope of the problem deserves stating: here, the label is the result. There is no independent ground truth — no reference translation produced by a professional interpreter, no second coder, no kappa coefficient. The whole fine structure of the taxonomy rests on the linguistic judgement of a single person, moreover an English speaker assessing Nepali at Advanced-Mid, a high but non-native level. This is the classic failure mode of translation-quality studies: the misleading metric here takes the form of unvalidated labelling.
The missing comparator is the human interpreter. We learn that 14.2 % of utterances are significantly altered. We have no idea what a professional interpreter, an ad hoc interpreter, Google Translate or a newer model would have produced on the same conversations. Yet the decision question is never "is ChatGPT-4o perfect?" but "is ChatGPT-4o better than what is available here and now?", and in the contexts the paper targets, the realistic alternative is often no assistance at all. Without a baseline the reader cannot answer, and the study does not claim otherwise. Add that the model tested, GPT-4o, was already a generation behind at posting: the figures date a configuration, not a technology.
Population bias, a tiny sample, and a counting inconsistency. Thirty adults recruited by convenience within five kilometres of a single hospital: generalization to other regions of Nepal, other dialects, older or less-schooled populations, or a clinical setting is neither tested nor established. Age, sex and literacy were collected but not incorporated into the analysis, so we do not know whether performance depends on who is speaking — which is the most obvious hypothesis to test. Finally an arithmetic detail: the abstract announces 329 deviation tags, but the counts and percentages in Table 3 imply a denominator of about 380 (65 tags equal 17.1 %, giving a total of 380). The gap does not change the ranking of failure modes, but a denominator inconsistent between abstract and table is the kind of detail peer review will correct — and a reminder that this is a preprint.
What it changes
For the research community. The paper shifts the evaluation question. As long as machine translation is measured on written text with similarity metrics, one optimizes a quantity — resemblance to a reference — that says nothing about risk. This work shows that under spontaneous-speech conditions the variable that matters is the detectability of the error by the user: a fluent, grammatically correct output that invents information is more dangerous than an obviously broken one, because the latter triggers a request to repeat and the former does not. That property is not captured at all by classic fidelity metrics, and it calls for work on flagging mechanisms — displayed back-translation, confidence scores, addition detection — rather than on additional BLEU points. The second open avenue is the asymmetry: if it holds up in a design that controls for input type, it means these tools understand better than they make themselves understood, which in a consultation amounts to amplifying the clinician's voice and degrading the patient's.
For clinicians. Nothing here justifies replacing an interpreter, and the authors are unambiguous about it. What the paper does provide is a list of warning signs usable immediately by anyone who ends up using a voice translator for lack of anything better: numbers (age, duration, dosage, frequency) are mishandled in 3.4 % of tags, which is enough to warrant systematically repeating and confirming every quantity; open-ended questions are the most degraded regime, which argues for short, closed questions; the patient-to-clinician direction is the most fragile, which makes it necessary to restate what you think you understood and have it validated. And above all: an answer that sounds right is not a faithful answer.
For patients and the public. The prevailing intuition is that AI makes visible errors — gibberish, a sentence that means nothing. This paper documents the opposite in a concrete, verifiable case: the system turned "I farm" into "I walk a little every day", and nobody in the conversation had any way of noticing. There is an equity asymmetry here that goes beyond the technical. The best-resourced languages — English, Mandarin, Spanish — get more reliable tools; lower-resource languages inherit tools whose appearance of quality is identical and whose fidelity is not. As these voice assistants spread into resource-limited health systems, it is that invisible difference, not average performance, that will determine who is heard correctly.
Further reading
The preprint: Accuracy and error patterns of ChatGPT-4o for real-time English–Nepali voice translation: A cross-sectional field evaluation in rural Nepal, medRxiv, DOI 10.64898/2026.08.25.26361303, posted 28 August 2026, under CC BY 4.0. Authors: Alyssa Mandich (School of Medicine, University of Washington), Shirsha Koirala (Northeast Ohio Medical University), Sophie Westen (Georgetown University), Saisha Adhikari (Association of State and Territorial Health Officials), Ashraya Acharya and Abha Shrestha (Department of Community Medicine, Kathmandu University School of Medical Sciences, Dhulikhel Hospital). The de-identified dataset of accuracy ratings and error codes is published as a supplement (S1 File); audio recordings and full transcripts are not publicly available and will be destroyed in July 2027, with access on justified request remaining possible through IRC-KUSMS. No competing interests declared. The project was supported by the University of Washington Global Health Immersion Program.