The evidence-to-record gap: what seventeen LLMs lose when they have to find the information in a real pregnancy chart
A team from Fudan University and its obstetrics and gynecology hospital built ObGynLongBench, an evaluation set of 1,500 decision points drawn from 976 real pregnancy records, where each question is anchored to a specific date with everything after it deleted. Tested on the same cases, seventeen language models answer correctly 58.7% to 79.1% of the time when the relevant evidence is handed to them, and only 41.3% to 68.7% when they must find it themselves in the chart — a gap of 7.6 to 22.7 points the authors call the "Evidence-to-EHR Gap". The finding is real and the statistics are careful, but the benchmark has no human reference performance, no control condition without the chart, and its labels were manufactured by the same model families it evaluates.
The context
For three years now, the reference measure for judging a language model in medicine has been the multiple-choice question: USMLE, MedQA, PubMedQA and their specialty variants. The format has one merit — it is objective and reproducible — and one flaw that grows as scores climb: the question arrives with its clinical vignette already written, filtered, summarised. What matters is right there, in three hundred words, with nothing else in the way. But that is not what a clinician does. A clinician opens a chart.
The electronic health record is long, redundant, badly ordered, and above all it contains an overwhelming majority of information irrelevant to today's decision. In this cohort, a single obstetric decision point is preceded on average by 479 laboratory results, 82 vital-sign measurements, 25 clinical notes, 6 ultrasound reports and 3 medication lines. The median duration of a followed pregnancy record is 184 days. End to end, that averages 30,800 tokens (the text units a model chops input into, roughly word fragments) with a standard deviation of 19,800 — far more than an exam vignette, far less than the context windows recent models advertise.
The question this team asks is therefore about a missing link: does a model that answers correctly when handed the evidence still answer correctly when it has to go looking for it? Earlier work explored the long record (MedAlign, MIMIC-Instr, TIMER), but in open-ended format, without anchoring to a traceable guideline or a strict temporal boundary. That double anchoring — an identified clinical rule, an enforced cut-off date — is what makes ObGynLongBench original.
The method
What is measured. Each case is a four-option question (A–D) about what to do, tied to a patient, a date in her pregnancy timeline, and a rule drawn from an obstetrics textbook or guideline. Everything after that date is removed from the chart given to the model; for inpatient cases the cut-off is set at admission and same-day diagnoses are erased automatically. The metric is a single one: accuracy. No AUC, no sensitivity, no calibration — the format does not lend itself to them, and that is a limitation in itself.
The three-scope design. This is the heart of the paper, and it is well built. The same cases are put to the same model three times, changing only what it is given to read. In Evidence-only mode, only the pre-located supporting items are provided. In Visit-level mode, all records from the same day. In History-level mode, the entire pre-decision chart. Comparing a model against itself across scopes cleanly separates two things: what the model knows about medicine, and what it can extract from a record. Cases are further graded into three levels: L1 (a single, locally visible piece of evidence), L2 (integration of several sources within the same visit), L3 (evidence distributed over time, typically more than seven days before the decision).
The data. 976 real pregnancy histories from a single tertiary hospital, in Chinese, spanning sixteen obstetric subspecialties (hypertensive disorders, gestational diabetes, intrahepatic cholestasis of pregnancy, group B streptococcus screening, multiple pregnancy and others). The 1,500 cases are not independent: 976 correspond to a patient's first decision point, 524 to a later one, with 240 patients contributing two. Collection period, age, parity and patient origin are not reported.
The models. Seventeen, none trained by the authors: four commercial "fast-tier" models (Claude-Haiku-4.5, GPT-5.4-mini, Gemini-3-Flash, DeepSeek-V4-Flash), eight long-window open-source models (Qwen, Gemma, Phi families) and five health-specialised ones (HealthGPT-Pro, Hulu-Med, Lingshu, MedGemma). Extended reasoning mode is disabled on the open and medical models, on by default for the commercial ones. Six chart-access strategies are then compared: direct reading, image rendering, two RAG variants (retrieving passages by similarity), rolling summary, and a tool-loop agent.
The construction itself relies on LLMs. Rule extraction from guidelines is done by GPT-5.3-Codex, L1/L2/L3 level assignment by Claude Opus 4.6, question and distractor writing by LLM, and validity verification by GPT-5.4-mini agents. Hold on to that point: it returns below.
The results
The gap exists, and it is wide. In Evidence-only, the best model is Gemini-3-Flash at 79.1%, followed by DeepSeek-V4-Flash at 78.8% and Gemma-4-26B-A4B at 78.2%. At History-level, Gemini-3-Flash stays on top but drops to 68.7%. Claude-Haiku-4.5 goes from 75.5 to 59.8 (−15.7 points), GPT-5.4-mini from 77.4 to 64.9 (−12.5), Qwen3.5-35B-A3B from 76.9 to 60.1 (−16.8). All seventeen models lose ground, without exception, from −7.6 (Phi-4-mini) to −22.7 (MedGemma-1.5-4B). The 95% confidence intervals, obtained by patient-clustered resampling, are about ±2.5 points: the gap is not noise.
Difficulty tracks how scattered the evidence is. For almost every model, L1 > L2 > L3. Gemini-3-Flash scores 66.0 / 71.5 / 65.0; Claude-Haiku-4.5 63.0 / 61.0 / 52.3; MedGemma-1.5-4B 45.8 / 43.5 / 29.7. That last figure deserves to be looked at squarely: with four options, chance sits at 25%.
Medical models are not better. The gap between the best generalist and the best health-specialised model widens from 4.2 points in Evidence-only to 7.4 points at History-level. Medical training, in other words, adds a little knowledge and nothing at all in the ability to read a chart.
Error propagates within a single patient. Across the 1,500 cases, models systematically do worse at a patient's second decision point than at her first — for all seventeen, from −1.86 to −10.88 points, mean −5.41, significant at p < 0.1 for twelve of them. This is the paper's most interesting safety result: an early misreading of the chart appears to contaminate later decisions for the same patient.
Among access strategies, only one genuinely helps. Against direct reading (59.6%), image rendering loses 6.9 points, static RAG and rolling summary are indistinguishable from noise, and only the active-search agent gains — +4.6 points, to 64.2%. Paired bootstrap confirms it: Image −6.91 ± 1.59 and Agent +4.53 ± 1.60 are the only two significant differences. Worth noting all the same: the agent consumes 35,786 tokens against 30,951 for direct reading, and compute budget is not equalised across strategies.
Clinical translation. Take 1,000 obstetric decision points of this kind. The best model, handed the evidence, gets 209 wrong. The same model, facing the full chart, gets 313 wrong: around a hundred decisions flip from right to wrong for the sole reason that the information had to be located. On L3 cases — where evidence is scattered over time, which is to say exactly what pregnancy follow-up looks like — MedGemma-1.5-4B gets 703 wrong, barely better than the 750 a coin toss would deliver. These numbers do not transfer to a delivery room: they are multiple-choice questions, not consultations. But they set a ceiling, and the ceiling is low.
What is good
The pre-decision information boundary is taken seriously. This is the paper's most careful engineering and the thing most rarely done properly. Everything subsequent is removed; so are admission-day diagnoses; an explicit scrub list strips earlier notes and prescriptions that would already state the chosen course of action; conversely, pre-cut-off labs and vitals are kept on principle, so the case does not become unsolvable. Without that discipline, you measure a model's ability to spot an answer already written in the chart — a close cousin of the circular leakage that ruins so many multimodal benchmarks.
The statistics are above the field average. The 1,500 cases are not independent, since 240 patients contribute several, and the authors handle it: confidence intervals by patient-clustered bootstrap (10,000 resamples, a drawn patient bringing all her cases), paired bootstrap to compare access strategies on the same cases, one-sided tests with per-model p-values published. Many medical AI benchmarks report a bare percentage; this one reports its uncertainty, and that uncertainty explicitly invalidates several micro-differences the main text nonetheless comments on a little too quickly.
An independent, quantified clinician audit. One hundred cases drawn at random from 1,500 were reviewed by three obstetrics residents who had not taken part in construction, with access to the full chart up to the cut-off. Aggregate: 96 reference answers judged correct, 1 incorrect, 3 uncertain; question quality judged clear and clinically relevant in 98 cases. Inter-rater agreement is reported via Gwet's AC1 (0.927 and 0.973) rather than a kappa, the right choice when annotations concentrate in a single category. Finally, the evaluation code is released under MIT licence, and every prompt appears in the appendix.
What is less good
The most elementary control condition is missing: the question without the chart. The benchmark's clinical rules come from public guidelines widely present in training corpora. No run is done with the stem and four options alone, with no patient data at all. So there is no way to know what share of the 79.1% — or indeed of the 68.7% — is obtained by simply reciting a guideline and eliminating distractors. This is a variant of the misleading metric: the chance level (25%) is never restated, no human performance is measured (the three residents audit the items, they do not answer them), and the numbers end up without a scale. If a substantial share of the score is reachable without opening the chart, the paper's central thesis weakens accordingly.
Labels and candidates come from the same industry. GPT-5.4-mini is the agent that verifies each item is valid and answerable from pre-cut-off evidence alone; it is also the second best model evaluated (77.4 in Evidence, 64.9 in History, best L3 score apart from Gemini). The retained items are, by construction, those it judges solvable. DeepSeek-V4-Flash is simultaneously the judge of evidence-utilisation scores and a model under evaluation — and, in the robustness matrix provided, it is the most atypical of the four judges (weighted kappa 0.73–0.77 against the others, versus 0.83–0.85 among them). The authors claim to have avoided circularity by choosing a closed format; they moved it from evaluation to label manufacture, which is not the same as removing it.
Part of the measured gap is probably a technical artefact, not a reasoning deficit. The two most spectacular drops — Hulu-Med-30A3 (−21.0) and MedGemma-1.5-4B (−22.7) — belong to the two models with the narrowest context windows (75,000 and 131,000 tokens), facing inputs of 30,800 ± 19,800 tokens. Three models in fact do better at Visit-level than at History-level, the classic signature of overflow. Yet no truncation rate, window-overflow rate or unparsable-response rate is published. On top of that comes an acknowledged but heavy asymmetry: extended reasoning on by default for commercial models, disabled for open and medical ones. Every cross-family comparison — including the conclusion that "medical models are not better" — is biased by it. Finally, note the population bias: a single Chinese tertiary hospital, in obstetrics, hence a caseload skewed toward high-risk pregnancy, with no demographic data published, and a release status for the 1,500 cases themselves that is never spelled out — only the code is declared under licence.
What it changes
For the research community, this paper usefully shifts the marker. It shows that a high score on a vignette-based multiple-choice test does not predict performance on a real chart, and it supplies a protocol — three scopes over the same cases — that cleanly separates medical knowledge from information extraction. It is reproducible elsewhere, in other specialties and other languages, and it should be. The active-search agent result also points somewhere concrete: better to let a model query the chart than to hope it reads it well in one pass. The within-patient error propagation result deserves dedicated work on its own.
For clinicians, nothing changes today, and the paper does not pretend otherwise: its ethics section states explicitly that ObGynLongBench is not a clinical decision support system. The operational message lies elsewhere. When a tool is presented to you with a benchmark score, the question to ask is: evaluated on written vignettes, or on the chart as it actually exists here? The gap documented here, ten to twenty points, is precisely the one that separates a demonstration from a deployment. It is the same pattern seen in conversational triage and in mid-course decision-making: performance collapses as soon as the situation stops being pre-digested.
For patients and the public, the lesson is simple and durable. A model that "knows" the guidelines is not a model that can read your chart. The two skills are distinct, the second is by far the less advanced, and it is the one that matters in pregnancy follow-up, where useful information is scattered across six months. A figure of 68.7% correct on a questionnaire says nothing about a tool's safety: it measures neither the ability to say "I don't know", nor the calibration of confidence — two properties this work, like most closed-format benchmarks, does not measure at all.
Going further
The preprint: ObGynLongBench: Revealing the Evidence-to-EHR Gap in Longitudinal EHR Decision-Making, arXiv:2609.07601, submitted 7 September 2026, DOI 10.48550/arXiv.2609.07601, under CC BY 4.0. The evaluation code and resources are announced on GitHub, under MIT licence; the release status of the 1,500 cases themselves is not specified.
Declared funding: National Key R&D Program of China (2023YFF1204800), the Shanghai municipal commission's "AI for Science" programme (2025-GZL-RGZN-BTBX-02028), National Natural Science Foundation of China (62406121), Shanghai Basic Research Program (25ZR1401034), with compute resources from Fudan University's CFFF platform. No conflict-of-interest statement appears in the manuscript; no author is affiliated with a laboratory producing any of the seventeen models evaluated.
On the same ground, the earlier long-record benchmarks cited by the authors: MedAlign, MIMIC-Instr and TIMER, all open-ended and without anchoring to a traceable guideline.