PrecepTron: a 32-billion-parameter judge calibrated to a physician panel, and what it reveals about clinical benchmarks

A team from Harvard Medical School, Beth Israel Deaconess Medical Center and Stanford releases PrecepTron, a 32-billion-parameter model fine-tuned to grade open-ended clinical LLM responses in place of a physician panel, together with GRAND-ROUNDS, a public set of 9,217 scores awarded by 11 physicians across seven already-published studies. The result that matters is not that PrecepTron reaches physician inter-rater agreement on three of five tasks, but that the three frontier proprietary models used as judges disagree with one another by up to 17 points on the same responses under the same rubric. And it is by using this judge that the authors discover GPT-5 scores well on 42% of New England Journal of Medicine cases from fewer than forty words of case text.

The context

Evaluating a language model in medicine long meant putting it through multiple-choice exams — the USMLE, question banks — where the answer is unique and grading is automatic. That era is over: models saturate these tests, and more importantly the tests resemble nothing a clinician actually does. The literature has therefore moved to open-ended responses: ask the model for a differential diagnosis, a management plan, a consultation note, and have physicians score its output against an explicit rubric.

That shift made benchmarks more credible and evaluation far more fragile. More credible because the comparator becomes a human on a real task. More fragile because the score now depends on who grades. Which experts do you recruit — generalists or subspecialists, average or expert physicians, academics or community practitioners? How many? Do they agree with each other? Studies published in JAMA, Nature Medicine or Science generally rest on panels of two to a dozen physicians, often from a single institution. Arjun Manrai, co-senior author here, summarised the problem in a 2026 Nature comment titled "Medical AI has a measurement problem". The FDA, for its part, opened a 2026 consultation on regulating generative-AI-enabled medical devices — meaning these scores will soon bear on regulatory decisions.

The fashionable answer is LLM-as-a-judge: hand the grading to another language model. It is fast, byte-for-byte reproducible, and already used to settle serious questions, such as whether medically specialised models beat general-purpose ones. Two obstacles. First, an automated judge can be poorly calibrated and carry biases of its own — a preference for longer responses, for AI-written text, for outputs from its own model family. Second, a proprietary judge runs at its vendor's: sending patient records to an external API is unacceptable to most hospitals. This paper attacks both fronts at once.

The method

The dataset. GRAND-ROUNDS (Graded Responses and Annotated Notes for Diagnostic Reasoning on UNstructured Data Sets) aggregates seven published studies in which physicians hand-scored LLM outputs: 9,217 physician scores over 5,250 response–rubric entries, produced by 160 clinicians and 9 AI models, graded by 11 physicians. Six tasks, five of which are used in the experiments: differential diagnosis on the NEJM Clinicopathologic Conferences (scored with the Bond score, a five-level ordinal scale — 0, 2, 3, 4, 5 — where 5 means the correct diagnosis appears in the differential); the Landmark Diagnostic Cases from Goh et al. 2024, scored on a 19-point rubric that separately credits candidate diagnoses, supporting and opposing evidence, the final diagnosis and proposed next steps; the Grey Matters Management Cases from Goh et al. 2025, on management reasoning; NEJM Healer, consultation-note quality under the 10-point R-IDEA rubric; and BIDMC ER, emergency-department diagnostic accuracy from the 2026 Science study. The sixth task — CPC testing plans — is released but excluded from the experiments because its scores cluster too heavily at the top of the scale.

The judges compared. Eight models in zero-shot mode, that is, with only the published rubric in the prompt: five open (Qwen3.5-9B, Qwen3-32B, Gemma-3-12B, Llama-3.1-8B, Mistral-3-8B) and three proprietary (Claude Opus 4.6, Gemini 3.1 Pro, GPT-5), plus an ensemble of the three proprietary judges whose score is their rounded mean.

The model. PrecepTron is a supervised fine-tune of Qwen3-32B via LoRA (low-rank adaptation): rather than moving all 32 billion weights, only small low-rank matrices grafted onto each attention and projection layer are trained, which makes the operation feasible on a single GPU and mechanically limits how far the model can drift. Rank 16, scaling factor 32, dropout 0.05, AdamW, peak learning rate 2 × 10⁻⁴, cosine schedule with 3% warm-up, effective batch size 4, 4,096-token sequences, bfloat16, three epochs, on a single 80 GB NVIDIA H100. A separate adapter is trained per rubric; the CPCs and BIDMC ER share one since they share the Bond score. Splits are made by case, never by response: 20% of cases for training on tasks with at least ten cases, only two cases for the two tasks with fewer — that is, 2 to 46 training cases depending on the task. One detail proves decisive: training entries are resampled by score tercile (low, mid, high) until the bins are balanced, capped at five repetitions, to stop the judge from settling into the most common score.

The measures. Primary endpoint, accuracy against the reference physician score: "within 1 point" on the Bond scale, "within 10%" of the normalised score elsewhere. Secondary endpoint, quadratic-weighted Cohen's κ — an agreement measure corrected for chance that penalises larger disagreements more heavily. Both are systematically compared to physician inter-rater agreement, computed on the subset of responses scored by at least two physicians. Ninety-five-percent confidence intervals by non-parametric bootstrap, 2,000 resamples, fixed seed. Two competing strategies that update no weights are tested on the same data: five-shot prompting and the GEPA prompt optimiser.

The results

Frontier judges contradict one another. On the same responses, under the same rubric, mean scores assigned by the three proprietary models differ by 17 percentage points on the Landmark Diagnostic Cases (Gemini 74%, GPT-5 59%, Claude 57%) and by 9 points on the CPCs (GPT-5 80%, Gemini 77%, Claude 71%, against 88% for physicians). And this disagreement is not a constant offset one could recalibrate away: Claude is the harshest on the CPCs (71% vs. 88% for physicians) and the most lenient on BIDMC ER (82% vs. 78%). No base model reaches physician inter-rater agreement across all five tasks. The authors checked on four of the five tasks whether a judge favours its own outputs: no systematic preference detected. The fifth, BIDMC ER, could not be tested — its text contains patient data, which does not leave the hospital.

Lightweight fine-tuning closes the gap. Trained on 2 to 46 cases, PrecepTron-32B gains 4 to 15 points over its base model on all five tasks: 82% → 92% on the CPCs, 46% → 61% on the Landmark cases. On that last and most demanding task, with its 19-point rubric, 93 scored responses drawn from two cases suffice to move κ from 0.62 to 0.80. The full table, accuracy and κ against the human baseline: CPCs, PrecepTron 92% (κ = 0.71) vs. 92% (κ = 0.68) for physicians; NEJM Healer, 81% (κ = 0.80) vs. 77% (κ = 0.83); BIDMC ER, 91% (κ = 0.60) vs. 88% (κ = 0.66); Landmark, 61% (κ = 0.80) vs. 67% (κ = 0.92); Grey Matters, 80% (κ = 0.78) vs. 95% (κ = 0.90). PrecepTron beats the three-model proprietary ensemble on two tasks (CPCs 92% vs. 83%, Landmark 61% vs. 44%) and matches it on the other three. Both prompt-based strategies fail: five-shot prompting drops NEJM Healer from 76% to 67%, GEPA degrades management reasoning from 72% to 63%.

Reproducing five studies. Applied to the original publications' responses with their original rubrics, training cases excluded, PrecepTron recovers the headline findings of five studies: GPT-4's diagnostic accuracy in Kanjee et al. (JAMA, 2023), the relative ordering of R-IDEA scores between chatbot, residents and attendings in Cabral et al. (JAMA Internal Medicine, 2024), GPT-4 performing comparably to or better than physicians with conventional resources in Goh et al. (2024 and 2025), and the rise in diagnostic accuracy from triage to admission in Brodeur et al. (Science, 2026).

The most disquieting result. With an automated grader, the authors can do what no human panel would fund: give GPT-5 progressively longer prefixes of each case, in steps of ten tokens, and score its response at each step. On the CPCs, GPT-5 places the correct diagnosis in its top-10 for 42% of cases from 50 tokens or fewer — roughly forty words — against 14% for a top-3 and 12% for a single diagnosis. For reference, a CPC case presentation runs about 1,600 tokens. On management cases, saturation is faster still: four of the five held-out questions reach 80% of the rubric maximum within 50 tokens, three of them within 10. All four Landmark Diagnostic Cases get there within 50 tokens. Gemma-4-31B behaves the same way qualitatively, needing somewhat more text. Memorisation does not explain it: the 50 CPCs used are from 2025 onward, after both models' pretraining cutoff, and the other two sets had never been published online. One example from Table 2 gives the measure of the thing: on "A 53-year-old man was evaluated in the cardiology clinic of this hospital because of severe left ventricular dysfunction with an apical aneurys", 30 tokens, GPT-5 offers Chagas cardiomyopathy as its second hypothesis — that was the diagnosis.

Concrete translation. This model touches no patient: it grades papers. The unit of damage is therefore not the false negative but the published conclusion. A 17-point gap in the mean score of a case set is exactly the span separating "the model beats physicians" from "the model falls short". If the choice of judge can move a mean that far, then part of the headline literature in health AI over the past three years rests on a parameter almost none of those studies report.

What is good

The negative result on frontier judges is the real contribution, and no recalibration recovers from it. Showing that three proprietary models grade the same papers 17 points apart would have been paper enough. The detail that matters is that these gaps are not internally consistent: Claude harsh here, lenient there. A constant bias is fixed with an offset; this one is not, it has to be measured per task. The demonstration is made with the means it required — eight judges, five tasks, two metrics one of which is chance-corrected, bootstrap intervals.

The ablations go looking for failure where it hides. Removing balanced resampling drops κ from 0.72 to 0.51 on NEJM Healer at comparable accuracy: the authors publish the evidence that their judge, without that precaution, collapses toward the majority score while still looking competent. They also test the most tempting shortcut — a judge that rewards anything resembling AI-written text — by retraining on human responses only, and find no shortcut. Hunting for one's own failure modes, then documenting them in the supplement, is rare.

The release is real, and the size makes local use possible. The 9,217 physician scores, the training code, the adapters and — for the first time — the case text of the Landmark Diagnostic Cases and Grey Matters are deposited on Hugging Face and GitHub. Human annotations from five influential studies become a reusable public good, which is precisely what lets others contest the rest of the paper. And 32 billion parameters fit on a single 80 GB GPU: a hospital can run the judge in-house. The argument is not theoretical — it is exactly because the BIDMC emergency-department text cannot leave the hospital that the three proprietary judges could not be tested there on self-preference.

What is less good

The "reproduction" of five studies is circular, and that is a biased comparator. Kanjee, Rodman, Crowe, Brodeur, Goh, Chen and Abdulnour are co-authors of this paper and of the studies it reproduces. PrecepTron is trained on 20% of the cases from those same corpora, with those same panels' scores, then "recovers" the results on the remaining 80%. That is an internal consistency check, not an independent replication: no study from another group, graded by a panel the authors did not train on, was used to test the apparatus. The authors say so plainly among their limitations — reproducing a study with PrecepTron reproduces that study's evaluation standard, not the underlying clinical judgment — but Figure 4 will read "reproduces five published studies", and that is the sentence that will travel. The same self-enclosed validation pattern was flagged here in a benchmark whose labels came from the model under evaluation.

The primary endpoint is permissive, and one metric contradicts the other where nobody looks. "Within 1 point" on a Bond scale with only five levels tolerates confusing a 4 with a 5 — that is, a differential that approaches the diagnosis with one that contains it. Above all: on BIDMC ER, PrecepTron posts 91% accuracy but its κ falls from 0.68 (base model) to 0.60, below the physicians' 0.66. Accuracy up and chance-corrected agreement down is the exact signature of a judge retreating to the frequent score of a skewed distribution — the very defect balanced resampling was meant to prevent, and the reason the sixth task was pulled from the experiments. The paper dispatches this in a clause ("declined modestly"). Add that the short-prefix finding is established with this judge, whose leniency is precisely what is in question; the authors flag one borderline case where PrecepTron is probably too generous, without a human re-read at scale to settle it.

What is frozen into the weights is a panel, and its composition is not neutral. The 11 grading physicians come overwhelmingly from the Boston–Stanford axis — Beth Israel Deaconess, Massachusetts General, Brigham and Women's, Harvard, Stanford. The paper presents as a feature the ability to "swap out" a panel, to replay a study with subspecialists or localise a national benchmark; but what ships is this panel, academic and American, whose latent preferences become the reference against which other models will be judged. One adapter per rubric, trained on 2 to 46 cases, and Step 2 of the protocol says it outright: validation on someone else's task does not transfer. The reusable artefact is therefore the recipe, not the judge. Finally, this preprint carries no funding statement and no conflict-of-interest declaration, the dataset licence is unspecified, and the BIDMC case text — the only task involving real emergency-department encounters — is not released, which makes that portion unverifiable from outside.

What it changes

For the research community, the operational message is twofold. First, the judge becomes a methods parameter to be declared like a reagent: model, version, prompts, calibration data, judge–physician agreement set beside physician–physician agreement. The five-step protocol given in a box is directly applicable and should become a peer-review requirement. Second, and this weighs more: if a model scores 42% correct on NEJM cases from forty words, then current rubrics largely reward pattern recognition — mapping an age, a sex and a symptom onto a prototypical presentation — and leave untested what clinical reasoning actually is: integrating new information, revising a hypothesis, resisting premature closure. Future benchmarks will have to vary information over time and credit revision, not only the end-of-case score.

For clinicians, nothing changes at the bedside today. The indirect effect deserves anticipating: if hospitals deploy LLMs clinically, they will need oversight at scale, and a local judge like this one is the likely form that oversight will take. Which makes the judge a new safety organ — and a new point of failure. A poorly calibrated judge does not produce a visible error: it produces a reassuring dashboard. The reading grid worth keeping fits in one question, facing any headline of the form "AI matches physicians": who graded, how many were they, and did they agree with each other? The same caution applies to clinical benchmarks scored by automated judges.

For patients and the public, one simple distinction. When a news item announces that a model "solved" complex cases, it is reporting a score, and that score was given by somebody — a small panel of physicians, or increasingly a machine. This paper shows that the somebody changes the result. It also shows that the exams used to crown these models can be partly solved by a first impression on a handful of words, which is a way of answering well without having reasoned. A model that looks like an excellent diagnostician on a benchmark has not shown it would change its mind when the patient does not fit the prototype. The same gap between vignette score and real behaviour has already been observed in conversational triage under real conditions.

Further reading

The paper: Scaling Clinical Judgment to Evaluate Medical AI, Thomas A. Buckley, Zahir Kanjee, Peter G. Brodeur, Byron Crowe, Anthony M. Pettinato, Aashna P. Shah, Adrian D. Haimovich, Liam G. McCoy, Daniel Restrepo, Ethan Goh, Jonathan H. Chen, Laura Zwaan, Katherine E. Goodman, Daniel J. Morgan, Raja-Elie E. Abdulnour, Adam Rodman and Arjun K. Manrai, arXiv preprint submitted 11 September 2026, category cs.AI, DOI 10.48550/arXiv.2609.12822. Not yet peer-reviewed.

Affiliations: Department of Biomedical Informatics, Harvard Medical School; Department of Medicine and Department of Emergency Medicine, Beth Israel Deaconess Medical Center; Stanford Division of Hospital Medicine, Division of Computational Medicine and Clinical Excellence Research Center, Stanford University; Division of Neurology, University of Alberta; Institute for Medical Engineering and Science, MIT; Department of Medicine, Massachusetts General Hospital; Institute of Medical Education Research, Erasmus Medical Center, Rotterdam; Department of Epidemiology and Public Health, University of Maryland School of Medicine and University of Maryland Institute for Health Computing; VA Maryland Healthcare System; Department of Pulmonary and Critical Care Medicine, Brigham and Women's Hospital. Adam Rodman and Arjun K. Manrai are co-senior authors.

Data and code: training and evaluation code, judge prompts and the few-shot and GEPA baselines at github.com/2v/PrecepTron; the PrecepTron-32B adapters in the Hugging Face collection; the GRAND-ROUNDS dataset, which for the first time includes the case text of the Landmark Diagnostic Cases and Grey Matters Management Cases; documentation at preceptron.net. The Beth Israel emergency-department case text contains protected health information and is not released. The dataset licence is not specified in the preprint.

Funding and conflicts of interest: the preprint carries no funding statement and no conflict-of-interest declaration. Seven of the seventeen authors are co-authors of the studies the paper sets out to reproduce (Kanjee et al. 2023, Cabral et al. 2024, Goh et al. 2024 and 2025, Brodeur et al. 2026).

Also on Tatakoto: a German clinical benchmark scored by an automated judge, circular leakage when the labels come from the model being evaluated, and the gap between vignettes and real triage.