MedQADE: an LLM grading medical answers matches physicians on agreement, never on caution
A team from the universities of Lübeck, Tübingen, Würzburg and Berlin's Charité builds MedQADE, the first open-response clinical benchmark in German: 3,800 questions drawn from the national medical exam, each scored both by physicians and by nine language models acting as graders. The best automated grader, Gemini 3 Flash, reaches agreement with the medical consensus (a kappa of 0.694) statistically indistinguishable from physicians' agreement with each other (0.709). But where a physician abstains when a question exceeds their expertise, frontier models issue a verdict in 100% of cases — surface-level agreement, without the clinical caution that normally comes with it.
The context
How do we know a language model is good at medicine? Since GPT-4, the standard answer runs through multiple-choice exams modeled on medical school tests (USMLE in the United States, MedQA). The problem is twofold. On one hand, the best models saturate these tests, whose scores no longer discriminate much. On the other, ticking the right box out of five has almost nothing to do with real clinical practice, where an answer must be produced from scratch. When the same questions are re-framed as open response (the AI must write, not choose), model performance drops — by about 39% in a recent study cited by the authors.
Open response is therefore more faithful, but it creates a bottleneck: a competent human is needed to judge each answer. Hence a practice that became common in 2025-2026, LLM-as-a-Judge: grading is delegated to another language model, meant to mimic expert judgment at scale and low cost. The question this paper asks is not "does the model answer well?" but "is the model that grades the answers a reliable judge?". Add a geographic blind spot: German, the language of tens of millions of patients, had no native open-response clinical benchmark. MedQADE aims to fill both gaps at once.
The method
The benchmark is derived from Ankizin, a peer-reviewed collection of flashcards heavily used by German students preparing for the second state medical examination. From the 44,185 cards of version 5, the authors keep only single-cloze items (one piece of information masked to be recovered), i.e. 26,598 items, then sample 3,800 in stratified fashion (80% with no exact lexical match to the reference answer, 20% with). Each item is an open question such as "What is the trigger of Guillain-Barré syndrome in two-thirds of cases?", with a short expected answer.
Five "student" models produce answers to all questions: Gemini 2.5 Flash, GPT-5 Nano, Gemma 3 27B, Gemma 3 4B and Qwen3-4B. These answers are then judged by two parallel instances. First, a panel of nine physicians practising in German hospitals (9.6 years of clinical experience on average, from 5 to 23), annotating through a dedicated interface. For each student answer, the physician picks among three labels: correct, incorrect, or abstain — the last to be selected explicitly if they feel they lack the expertise to decide. They also rate the perceived difficulty of the question. A tenth physician acts as tiebreaker. Then, nine language models act as automated graders: the five students plus four extra models (Gemini 3 Flash, Gemini 2.5 Flash-Lite, GPT-5.4 Nano, GPT-5.4 Mini), so as to also measure self-enhancement bias.
The methodological core is how a grader is compared to the consensus. Comparing a judge directly to the full set of physicians mechanically inflates its score. The authors therefore build a human ceiling by leave-one-out validation: for each physician, the consensus of the other eight is formed, and the held-out physician's agreement with that consensus is measured via Cohen's kappa (an agreement measure corrected for chance: 1 = perfect agreement, 0 = chance-level). This yields the bar an average human reaches against their peers, with a confidence interval estimated by bootstrap resampling. An automated grader is deemed "clinician-level" only if its interval overlaps the human ceiling's.
The results
Among physicians, mean agreement is kappa 0.61 — "substantial" in the usual sense. The human ceiling computed by leave-one-out stands at kappa 0.709 (95% confidence interval: 0.667 to 0.746). The best automated grader, Gemini 3 Flash, reaches 0.694 (0.619 to 0.754), a gap of only 0.016 below the ceiling, with widely overlapping intervals: statistically, it cannot be distinguished from an average human physician. Two other models (GPT-5.4 Mini at 0.616, Gemini 2.5 Flash-Lite at 0.602) still overlap the ceiling; all others fall below, down to Gemma 3 4B at 0.327. Combining several graders into a "majority vote" improves nothing: Gemini 3 Flash alone remains the best.
Yet the most telling result is not that agreement, but what accompanies it — or rather what is missing. Physicians scale their abstention with difficulty: the harder a question, the more often they tick "I cannot decide". Frontier models, by contrast, assign a definitive label (correct or incorrect) in 100% of cases, regardless of difficulty. Only the small open models abstain at all (Gemma 3 4B: 6.41%, Qwen3-4B: 4.23%), and even then an order of magnitude below humans. The authors name this deficit an absence of clinical metacognition: the ability to recognize the limits of one's own knowledge.
Measured biases add to the picture. Gemma 3 4B overscores its own answers by 16.29 percentage points against its peers' consensus; within an architecture family, GPT-5.4 Mini grants +6.63 points to GPT-5 Nano, and Gemma 3 4B favors its 27B sibling by +11.54 points. As for raw student performance, under a pass threshold set at 60%, only the two commercial models (Gemini 2.5 Flash and GPT-5 Nano) would pass; the small models collapse to 16-17% correct on items without lexical match. In concrete terms: this paper does not test a diagnostic tool, it tests the reliability of the grade. And if an AI is used to declare that another medical AI is "correct" — which validation dossiers already do — its overconfidence becomes a safety problem, not a curiosity.
What is good
A real gap filled with rigor. MedQADE is the first open-response clinical benchmark in German, built on a peer-reviewed corpus rather than on questions generated on the fly. Above all, the leave-one-out human ceiling avoids the classic mistake of comparing an automated judge to the full set of experts (which mechanically favors it). Bootstrap confidence intervals and the exclusion of cases without a clear majority (2% of the data) show methodological care above the field's average.
A specific, useful finding. Quantifying graders' lack of caution — 100% definitive verdicts versus human abstention rising with difficulty — turns a vague intuition ("the AI is too sure of itself") into an actionable number. This is exactly the kind of measure missing whenever the reliability of automated evaluations is debated.
Transparency about its own blind spots. The authors report self-enhancement and intra-family biases, acknowledge wide confidence intervals, possible data contamination and a small panel. The dataset and code are announced as public upon publication, and the study is publicly funded (German federal research ministry, SEPE project), with a declared absence of competing interests.
What is less good
A possible data contamination, admitted but not controlled for. Ankizin is public and widely circulated; frontier models have likely seen it during training. This data leakage — when test material has, directly or not, served for training — inflates both student performance and judge agreement, without any way to separate it from genuine competence. The authors acknowledge it but provide no overlap analysis.
A headline stronger than the figures. "Clinician-level" rests on a 200-item subset and on wide confidence intervals (about 0.13 kappa in span) that the authors themselves call "to be interpreted with caution". The 0.016 gap below the ceiling is within statistical noise, and the human ceiling itself is only 0.709 — substantial agreement, not absolute. This is a classic misleading metric: overlapping intervals are not a demonstrated equality. Worse, physicians' agreement on question difficulty is weak (ordinal alpha of 0.199), which undermines the difficulty stratification on which the flagship abstention finding precisely rests.
A partial comparator and a provenance to verify. The graders tested are "lite" variants or small open models, not the most powerful configurations; the human panel is small (nine physicians plus a tiebreaker) and a single reference answer, provided "for orientation", frames the judgment. Finally, the preprint is dated 1 July 2026 and cites models with unusual names (GPT-5.4, Gemini 3 Flash): provenance and reproducibility remain to be confirmed as long as the announced data and code are not actually deposited.
What it changes
For the research community, MedQADE usefully shifts attention: evaluating no longer only the models, but the evaluators. The idea of using abstention as a measure of caution, and the leave-one-out human ceiling method, are transferable to other languages and specialties. The work also provides an infrastructure block for a language until now lacking a native open-response clinical benchmark.
For clinicians, nothing changes at the bedside. The practical message is defensive: beware of validation dossiers that certify a medical AI as "correct" on the strength of automated grading. A grader that never abstains can reach the average agreement of a physician while masking its incompetence on hard cases — precisely where error is costly.
For patients and the general public, the hook "an AI reaches physician level" is exactly the one to distrust. Here it is a grading task on exam flashcards, in retrospective conditions, and the real finding is not the machine's competence but its confidence without humility. The gap between an agreement score and a trustworthy clinical judgment — the crux of medical AI — remains wide open.
To go further
The preprint is available on arXiv (10.48550/arXiv.2607.01103). Work funded by the German federal research ministry (BMFTR, SEPE project, grant 16SV8958); the authors declare no competing interests, one co-author being affiliated with a private company (Genie Enterprise Deutschland GmbH). The dataset and code are announced as public upon publication. On other uses of LLMs in the clinic and their limits, see our decryptages on an LLM patient simulator, on the clinical safety of frontier LLMs, and on the faithfulness of LLMs in public health.
Editorial transparency: French version written and signed by the Tatakoto editorial team based on a reading of the preprint. English, Spanish and Chinese translations produced with AI assistance and reviewed.