Forty percent of a diagnosis column filled in by default: an account-by-account audit of a Chinese cognitive screening programme, and twenty-four ways to train a model on it

Fourteen Chinese authors opened the cognitive-status column of a community screening programme running in routine service across mainland China — 112,023 rows written by 2,597 data-entry accounts — and found that 181 accounts, each having entered at least 100 diagnoses without ever recording a single impairment, produced 45,315 of those rows, or 40.5% of the column. They then trained 24 models on the same patients under a specialist reference standard: neither fine-tuning, nor preference optimisation, nor reinforcement learning on a few hundred cases beat the 21-variable logistic regression already in production (AUROC 0.926). The one arm that does beat it — a Qwen3-4B running on a consumer GPU, at 0.940 — saw no physician label at all: it was distilled from 943 verdicts of a frontier model queried on the programme's own unlabelled records.

The context

Routine service databases — the ones a nurse, a social worker or a physician fills in at every contact, in the normal course of care — have become the preferred label source for clinical models. They are enormous and free. The problem, stated bluntly by the authors, is that the process that writes those labels is almost never audited before anyone trains on it. The label noise literature knows how to handle random or class-conditional noise; it is poorly equipped against noise concentrated in a handful of data-entry accounts, which no increase in sample size averages away.

The setting is a community cognitive-screening programme in routine use across mainland China, whose data are held by the Beijing Medical Award Foundation. Each resident takes two instruments: the AD-8, eight questions put to a close contact (an informant interview: it is the family, not the patient, who answers), and the HKBC (Hong Kong Brief Cognitive Test), a CSI-D-derived test with nine cognitive items plus immediate recall. A 21-variable logistic regression — the 8 AD-8 items, the 9 CSI-D items, immediate recall, age, sex, education — flags residents for referral at a probability threshold of 0.10.

This database holds two assets of opposite nature. One large and cheap: 112,023 cognitive-status rows, covering 105,893 assessed individuals — 103,688 recorded normal, 2,197 unspecified cognitive impairment, 1 mild cognitive impairment and 7 dementias. One small and expensive: the 672 individuals whose status was recorded by a titled physician (attending rank or above). The paper's question is what each way of using these two assets buys, and at what price.

The method

The audit. For each of the 2,597 operator accounts the authors count two things only: diagnoses entered, and the proportion recording any impairment. The detection rule flags accounts with at least 100 diagnoses and zero impairments. The crucial methodological move: they test a competing explanation. If those zeros came from a bulk process (an automated import), they should cluster in write timestamps. And indeed they do — the densest single second in the table contains 26,694 rows — but the effect hits every operator class, and titled physicians' rows are more timestamp-concentrated than untitled accounts' rows. Conclusion: the write-time column is a one-off database migration artefact, not an auto-fill signature. It is disqualified; the operator axis stands. No temporal claim is made anywhere in the study.

The 24 arms. They span the supervision choices a real programme faces. The 21-variable regression refitted on physician labels is the bar. Four local language models — Qwen2.5-1.5B and 3B, Qwen3-4B and 8B, loaded in 4-bit, LoRA-trained on a single 8 GB consumer GPU — appear zero-shot, with an explicit chain-of-thought pass, fine-tuned on physician labels, fine-tuned on routine labels with and without "decontamination" (dropping rows entered by accounts with ≥ 200 diagnoses and zero impairments: 54 accounts, 27,616 rows), or fine-tuned on a proxy task — the HKBC score band — which deliberately occupies a different label space from the clinical verdict. Added to these are DPO (direct preference optimisation) and GRPO (group relative policy optimisation) variants, the pre-registered two-stage "proxy then reinforcement" recipe at three false-negative penalties, a proprietary frontier model (gpt-5.6-pro) queried zero-shot, and finally distillation of that frontier model into the regression and into the local 4B.

Distillation, costed. The programme's unlabelled pool is reduced to 42,823 rows, which collapse to 15,127 distinct rendered patterns. Label-free stratified sampling (AD-8 endorsements plus CSI-D errors, nine bands) with stored inclusion probabilities yields the 943 patterns actually labelled by the frontier model. The 4B student is trained on them with weighted cross-entropy (Horvitz–Thompson weights, mean 1, maximum 44.8), in seven minutes. Total supervision cost: 943 calls at roughly 12 seconds each, about three hours of teacher time. Not a single physician label enters this arm's training.

The evaluation. One five-fold split, stratified on outcome and grouped on institution clusters resolved through the platform's team registry, so that no cluster spans a fold boundary — checked programmatically. The authors note in passing that an earlier lineage of the project compared raw institution name strings and leaked; the released folds are re-cut. Every arm is scored out-of-fold. One entire institution (Peking University Sixth Hospital) is held out, leaving a development panel of 642 individuals at 38 institutions, 259 of them impaired (40.3%). Paired contrasts use a 2,000-draw cluster bootstrap. Metrics: AUROC, expected calibration error (ECE), calibration slope, saturation, and net benefit at the deployed threshold of 0.10, reported per 100 individuals against the best trivial policy.

The results

The audit. The 181 flagged accounts wrote 45,315 rows, 40.5% of the column; six accounts each entered over 1,000 diagnoses without one impairment, the largest at 3,960. Recorded impairment falls monotonically with account volume: 15.7% for accounts with 1–9 rows, 3.4% at 10–49, 2.1% at 100–499, 0.7% at 500–999. Dropping the flagged accounts raises the programme's observed impairment rate from 2.60% to 4.37% — the apparent prevalence of a national screening programme nearly doubles through the exclusion of a few hundred data-entry accounts alone.

The ranking of arms. Under the specialist standard, the incumbent regression scores an AUROC of 0.926 [0.873–0.956]. No locally trained arm beats it. Fine-tuning on physician labels tops out at 0.924 (4B as well as 1.5B); DPO falls to 0.881, GRPO to 0.789 — that is −0.043 and −0.135 against their supervised sibling, at identical start, folds and hyperparameters. The pre-registered two-stage recipe is worse than its single-stage contaminated baseline (−0.030; 95% CI −0.077 to −0.004), and worse still than the ablation that swaps reinforcement for cross-entropy (−0.057). The arm trained on the contaminated routine corpus (0.899) is indistinguishable from the same untrained model (0.890; Δ +0.009, −0.014 to +0.024): the contaminated supervision bought nothing. The arm trained only on the proxy task collapses to 0.792, that is 0.098 below its own zero-shot base.

Chain-of-thought degrades. Asking the model to reason explicitly before answering reduces discrimination at every size: −0.072 at 1.5B, −0.080 at 3B, −0.041 at 4B, and −0.012 (not significant) at 8B. The penalty shrinks with scale but never becomes a gain.

The frontier model and its student. gpt-5.6-pro zero-shot reaches 0.932, statistically indistinguishable from the regression (+0.007, n.s.): as a predictor, it adds nothing. As a labelling instrument, it is another matter. The distilled 4B reaches 0.940 [0.898–0.963], above the regression (+0.014; 0.004 to 0.031) and above its own teacher (+0.008; 0.001 to 0.017). The label budget saturates very early: 0.937 with just 50 teacher labels, 0.940 by 200, 0.940 at the full 943.

Calibration separates better than discrimination. Only three arms are well calibrated — the regression (ECE 0.039), the frontier model (0.041) and the distilled 4B (0.045) — and they are exactly the three whose fixed-threshold decision value holds up. The 4B fine-tuned on physician labels shows an ECE of 0.102, a slope of 0.40 and 68% saturated outputs. The 1.5B, 3B and 8B zero-shot arms flag 100% of the panel at the deployed threshold: they collapse onto the trivial refer-everyone policy. The DPO arm saturates 99% of its outputs. In decision terms: per 100 individuals assessed, the distilled arm delivers a net benefit of +3.13 against the best trivial policy, the teacher +2.94 — but the authors warn that in the companion paper, where these contrasts come with intervals and multiplicity correction, no contrast between deployable arms excludes zero.

The held-out site reverses everything. On the 58 individuals of the held-out hospital, the ranking reorders completely: zero-shot 3B 0.925, zero-shot 4B 0.894, 8B 0.875, frozen regression 0.869, frontier model 0.851, distilled 4B 0.843, zero-shot 1.5B 0.808. The panel's winner finishes second to last. The authors say it themselves: no arm is separable from any other there, and the check is reported for transparency, not for selection.

What is good

The competing-explanation test is what separates an audit from an anecdote. It would have been easy to publish "40% of the column comes from suspicious accounts" and stop. The authors explicitly constructed the rival hypothesis — bulk filling identifiable in timestamps — tested it, and refuted it by showing that the temporal concentration also hits titled physicians. The consequence is operational, not cosmetic: had the timestamp hypothesis survived, the remedy would have been "drop bulk-written rows"; the correct remedy is "drop the flagged accounts' rows, and stop trusting the date column". They go so far as to forbid themselves any temporal claim in the rest of the paper, and to disclose that an earlier version of their own pipeline leaked on institution name strings.

The pre-registered results are published in the negative. The announced recipe — abundant proxy first, reinforcement on specialist cases second — did not work, and was in fact worse than the contaminated one-stage version. The authors report it that way, with three quantified contrasts and three penalty values, rather than reframing the objective after the fact. They additionally replicate the SFT/DPO/GRPO comparison at 4B to show that the failure of the reinforcement objectives is not a small-model artefact, and honestly note that in this two-action setting GRPO reduces to a vanilla policy gradient and DPO to a logistic margin loss — so that nothing here says anything about reinforcement learning on long-horizon reasoning tasks.

The distillation is costed all the way down. Call budget (943), throughput (median 12.1 s, zero unparseable responses on the first 472-call tranche), teacher noise measured by re-labelling 100 patterns (24% of repeats identical, 72% within ±0.05, mean absolute range 0.053), a label-budget efficiency curve, and a train–test overlap sensitivity analysis (16 of the 944 patterns match a panel pattern, covering 18 of 642 individuals; excluding them moves AUROC from 0.940 to 0.940). This is the level of accounting missing from nearly every clinical distillation paper.

What is less good

The audit, which is the heart of the paper, is entirely descriptive. The monotone gradient from 15.7% to 0.7% comes with no test, no interval, no p-value. The "refutation" of the timestamp hypothesis rests on a qualitative comparison — titled physicians' rows are "more concentrated" — with no quantification of that gap. The authors themselves call their audit a lower bound, since it detects accounts that never record a positive and misses any subtler default-filling behaviour. The conclusion is probably right; it is presented without the apparatus that would let a reader check it.

The panel is enriched to 40.3% impairment while the programme records 2.6%, and there is no patient characteristics table at all. The authors flag the first point and explicitly refuse to claim population performance — to their credit. But every pooled AUROC in the paper lives in that enriched world, and the misleading metric failure mode is right there: an AUROC of 0.940 on a cohort with 40% positives says nothing about what the model would do at threshold 0.10 in the real screening population. More awkwardly, the age, sex and education of the 672 individuals appear nowhere as values: they figure only as input variables. It is therefore impossible to assess population bias — which age range, what proportion of women, what schooling level, in a country where formal education among older cohorts varies enormously by province. The same blind spot affects the evaluation: the effective sample size is estimated near 158 out of 642 individuals, with an outcome intraclass correlation of 0.360; several null results are therefore indeterminate rather than established equalities, as the authors acknowledge.

The word "pre-registered" appears four times and is never accompanied by an identifier. No registry, no OSF number, no protocol DOI, no filing date. In a paper whose central argument is that you must audit data provenance before using it, the absence of verifiable provenance for its own protocol is a costly irony. Three weaknesses of the same family follow. Data and code are announced as "frozen, script-regenerable artefacts" but remain available on request from the corresponding author — nothing is publicly deposited. The institution used as the held-out site is exactly the affiliation of the corresponding author and of everyone thanked for data curation; this is not external validation, and the authors do not present it as such. Finally, the 1.5B and 3B arms are Qwen2.5 while the 4B and 8B are Qwen3, so that every cross-size comparison confounds scale with model generation — including the paper's most quotable conclusion, that the chain-of-thought penalty shrinks as the model grows.

What it changes

For the research community. The transferable result is not the 0.940 score, it is the two-column procedure. Operator identifier, volume, positivity: those three fields already exist in nearly every service database, and suffice to discover that 40% of an outcome column is default-filled, before a single model is trained. The second lesson is more uncomfortable for the current fashion: on a few hundred specialist cases, a well-calibrated 21-variable logistic regression is beaten by no fine-tuning, no preference optimisation, no reinforcement. The ceiling is set by the instruments and the reference standard, not by model capacity. It is the same verdict as in our decryption of severe acute pancreatitis, where random forest holds its own against deep learning, and it keeps holding.

For clinicians and programme managers. The number to remember is 2.60% versus 4.37%. The apparent prevalence of a national screening programme depended, for roughly half of it, on the data-entry behaviour of a few hundred accounts. Every management indicator built on that column — coverage, referral rate, screening yield — was wrong in the same proportion, independently of any artificial-intelligence question. The second practical point concerns calibration: only three arms produced probabilities usable at the programme's 0.10 threshold, and several models otherwise strong on discrimination flagged everybody. An AUROC says nothing about what happens at a given threshold, as our reading of ICU mortality scores and their decision curves already showed.

For patients and the public. Out of 105,893 people assessed by this programme, the database recorded 1 mild cognitive impairment and 7 dementias. That is not an epidemiological reality: it is the trace of a column filled in by default. The consequence is not a bad diagnosis made by a machine, but something more ordinary and more widespread — people genuinely screened whose referral was never recorded, and a public-health statistic that understates need. The paper's final paradox deserves to be stated as is: the study's best model never saw a single physician label, and its superiority over its own teacher can be explained just as well by its averaging over teacher noise as by the fact that the specialist reference standard is itself a noisy routine record, single reader, no adjudication. The two mechanisms are indistinguishable here, and the authors do not claim to separate them.

Further reading

The preprint: Default-filled outcome labels in a deployed cognitive-screening programme: an operator-level audit and the construction of twenty-four language-model arms, medRxiv, DOI 10.64898/2026.08.28.26361585, posted 2 September 2026, under a CC BY 4.0 licence. Authors: Jun Ji (Qingdao University, submitting author), Zhigang Sun, Xiaofang Ying, Jinyu Hao, Zhiqiang Fu, Dongmei Shi, Xiaoming Kong, Yijie Xu, Xiaojuan Zhang, Xiaoli Du, Zhiyan Zhang, Xinhong Liu, Ping Lin, and Huali Wang (Peking University Sixth Hospital, corresponding author) — fourteen authors spread across mental health centres and community health centres in ten Chinese provinces. Ethics approval: Ethics Committee of the Medical College of Qingdao University, no. QDU-HEC-2025431, 22 May 2025. No specific funding, no competing interests declared. The de-identified dataset, analysis code and derived artefacts are announced as available on request from the corresponding author; identifiable source records are held by the Beijing Medical Award Foundation and cannot be redistributed. The authors declare having used Claude (Anthropic) extensively and under their direction to write and revise the analysis and training scripts, to build the verification tooling that cross-checks manuscript numbers against the frozen result files, and to draft the text. A companion paper, announced but not yet identifiable, reuses the frozen predictions of this study to quantify the effect of evaluation-design choices relative to model choice. Instruments cited: AD-8 (Galvin et al., Neurology 2005) and the Hong Kong Brief Cognitive Test, derived from the CSI-D.