ChatGPT Health for triage: more conversation does not make the advice safer

Alvira Tyagi, Girish Nadkarni, Bilal Naved and colleagues (Icahn School of Medicine at Mount Sinai and partners) put ChatGPT Health through 255 triage cases — clinician-written vignettes, real emergency-department visits and real nurse-line calls — to test whether talking to the AI helps steer a patient to the right level of care. The result: exact agreement with human standards stays around 50%, and, crucially, when the AI diverges from the nurse standard it under-estimates urgency in nearly seven cases out of ten, with a dangerous under-triage roughly one time in ten. The lesson is not a score: it is a direction of error — the conversational fluency of a consumer assistant masks a systematic bias toward false reassurance, and talking longer does not fix it, it makes it worse.

The context

More and more people describe their symptoms to a conversational assistant before deciding whether to stay home, call a doctor or head to the emergency room. That seemingly trivial act is in fact a triage task: turning incomplete, patient-reported information into a level of urgency. Clinical triage is a decision under uncertainty that balances two opposing risks — over-triage, which sends people needlessly to the ER and clogs the system, and under-triage, which delays needed care. The direction of the error therefore matters as much as its frequency.

ChatGPT Health is a consumer health assistant built on a large language model (LLM) — a system trained to produce natural-language text from a prompt. The argument from its defenders ran as follows: past evaluations, often based on a single fixed question, would under-estimate systems designed for dialogue. Since a triage nurse asks successive questions to refine a decision, a model able to do the same should steer better. It is precisely this assumption — more conversation equals better triage — that this work puts to the test.

The method

The authors run a retrospective, cross-sectional evaluation of ChatGPT Health on 255 cases from three sources: 39 standardized vignettes written by clinicians (spanning 21 clinical domains), 76 real visits to a community emergency department in Maryland, and 140 real calls to a nurse triage line at a large US hospital system. Each case is presented under two regimes. In the "natural" mode (single-turn), the model receives a single sentence describing the symptom and then produces its recommendation — the typical use of a hurried patient. In the "multi-turn" mode, the model is instructed to ask one question at a time, like a nurse, before concluding; the graders answer from the record, and "I don't know" when information is missing.

Each recommendation is graded on a four-level ordinal scale: A (stay home), B (see a provider beyond 24 h), C (seek care within 24 h) and D (go to the ER), D being the most urgent. Two references serve as the standard: physician adjudication and nurse triage based on the Schmitt-Thompson protocols, a recognized standard for telephone triage. The team measures exact agreement, the weighted Cohen's kappa (a chance-corrected agreement index: 0 = chance, 1 = perfect), the mean ordinal distance (the number of steps apart), and above all the direction of the error — under- or over-triage — with two critical categories: safety-critical under-triage (at least two levels below the reference, the dangerous mode) and utilization-critical over-triage (at least two levels above, the costly mode). All evaluations were run between March 5 and May 21, 2026 on the standard consumer app, with no custom system prompt, history cleared between cases.

The results

Exact agreement is mediocre and stable. Against nurse triage, ChatGPT Health gets it exactly right in 52.9% of cases in natural mode and 55.7% in multi-turn; against physician adjudication, 54.1% and 48.2%. The weighted kappas (0.36 to 0.44) reflect only fair-to-moderate agreement. In other words, the model matches the human standard about half the time — and switching to dialogue does not change that half.

The real signal is elsewhere. When the AI diverges from nurse triage, the gap leans heavily to one side: 70.8% of disagreements are under-triage in natural mode (P < 0.0001) and 69.0% in multi-turn (P < 0.0001). The model under-estimates urgency far more often than it over-estimates it. Safety-critical under-triage — advising someone to stay home or wait where the reference would send them to the ER — affects 10.2% of cases in natural mode and 8.6% in multi-turn against the nurse standard. In clinical terms: out of 1,000 people consulting this way, about a hundred would receive dangerously reassuring advice. And this severe under-triage worsens moving from synthetic vignettes (5.1%) to real cases (10.5% in the ED, 17.9% on the nurse line): the more a case resembles real life, the more frequent the dangerous error.

Dialogue does not help. Exact agreement does not differ significantly between the two regimes (McNemar's test, P = 0.43 against nurse triage), and against the physician standard, multi-turn even tends to drift further away. Worse, in multi-turn, the longer the conversation, the further the recommendation strays from the right one (Spearman correlation ρ = 0.162, P = 0.010) — an association absent in natural mode. The authors honestly note that hard cases naturally invite more questions, which blurs any causal reading, but the finding holds: lengthening the conversation does not bring it closer to the right disposition. Finally, performance varies by complaint: respiratory (75.0% agreement) and neurological (70.0%) presentations fare better than dermatologic (45.2%), gastrointestinal (48.8%) or unclassifiable (38.5%) ones.

What is good

A protocol that mixes vignettes and real patients. Rather than settle for laboratory vignettes, the study confronts the model with 216 real cases (ED plus nurse line) adjudicated by clinicians. That is what reveals the most worrying fact: safety-critical under-triage rises from 5.1% on vignettes to 17.9% on real calls. A purely synthetic test would have falsely reassured.

Measuring the direction and severity of error, not just the average. Here lies the conceptual contribution: separating aggregate agreement (around 50%, unremarkable) from the structure of the error (a 70% under-triage bias, dangerous). By reporting ordinal distance, direction and critical categories, the authors show that the same "50% agreement" can hide a population of errors that are safe or deadly depending on their sign.

A head-on test of the "more conversation = better" assumption. The study does more than note a weakness: it directly refutes the classic argument that single-turn evaluations are unfair to systems designed for dialogue. By showing that multi-turn does not improve — and that long conversations worsen accuracy — it answers a precise, recurring objection with an experiment designed for exactly that.

What is less good

The misleading metric, here turned inside out but still latent. The article fights this trap itself, but it is a reminder of how widespread it is: a press release announcing "ChatGPT Health agrees with nurses half the time" would be technically true and clinically dangerous, since the failing half leans toward false reassurance. The takeaway: overall agreement says nothing about safety until you have looked at the direction of the errors.

A gold standard that is not quite gold, and no prospective test. The two human references themselves agree in only 59.6% of cases, and nurse triage tends to be more cautious than physician adjudication — so the "right answer" is fuzzy. Above all, the evaluation is retrospective: it relies on vignette texts and chart notes, not on real live conversations with real patients who answer awkwardly, forget details or panic. Real-world performance could be even worse — or simply different. It is an acknowledged measurement bias, but a real one.

A single system, trainee graders, and a conflict of interest to note. The work tests only one product (ChatGPT Health) at one moment (March–May 2026); a fast-moving model could behave differently tomorrow. The graders coding the AI's answers are mostly medical students and one doctoral fellow, with only fair inter-rater reliability for some (an outlier grader had to be excluded in a sensitivity analysis, without changing the conclusions). Finally, one senior author declares a financial interest in Clearstep, a digital triage company — disclosed and managed, but worth keeping in mind since the conclusion, favorable to deterministic guardrails over "pure generative", points in the direction of such products.

What it changes

For the research community, the methodological message is clear: evaluating a consumer health assistant on its aggregate agreement rate alone is no longer enough. One must measure distance from the reference, the direction of the error, the acuity levels where it concentrates, and test in real-world conditions, not only on vignettes. The study provides a template — and public code — to do so.

For clinicians, an immediate practical consequence: add to the history-taking the question "did you consult an AI before coming in?". A patient arriving with model-generated false reassurance may have delayed or be downplaying symptoms; sometimes they need to be "de-anchored" from wrong advice. Since under-triage concentrates on high-acuity cases (levels C-D), the model fails precisely where the error costs the most.

For patients and the public, the lesson is sober: a conversational assistant's fluency is no guarantee of safety, and a longer dialogue does not make it more reliable. Faced with an acute symptom — chest pain, difficulty breathing, neurological signs — the tool may reassure wrongly. These platforms need explicit warnings and deterministic guardrails, and their use as a de facto triage tool would, the authors argue, warrant regulatory oversight as software as a medical device.

Going further

The preprint is available on medRxiv (DOI 10.64898/2026.07.21.26358588), with no external funding, under a CC BY-NC-ND 4.0 license; the analysis code is on GitHub and the data on Zenodo. For context, see the structured evaluation of ChatGPT Health by Ramaswamy et al. (Nature Medicine, 2026) and the Schmitt-Thompson nurse triage protocols.