Three minutes of video against the false positive paradox: 34 radiologists, and acceptance of false AI flags cut in half
Thirty-four radiologists from ten countries were randomized into two arms, then read the same 20 brain MR angiograms, each carrying an aneurysm flag produced by a commercial tool — 18 of those flags being false. The arm that had watched a three-minute video on the false positive paradox accepted 12.7% of false flags versus 22.5% in the control arm, and recommended a follow-up examination in 39.2% of cases versus 54.9%. The most telling number in the paper lies elsewhere: before any reading, the 34 participants estimated on average a 62.9% probability that an AI flag corresponds to a true aneurysm, whereas simulation from published performance figures gives 15.4%.
The context
A diagnostic test can be excellent and still be wrong most of the time when it says "yes". This is the false positive paradox: in a population where the disease is rare, even high sensitivity (the proportion of diseased cases correctly detected) and specificity (the proportion of non-diseased cases correctly ruled out) produce more false alarms than true ones. The quantity that matters then is neither of these but the positive predictive value (PPV): among all flagged cases, what fraction is genuinely positive. PPV depends on prevalence, not only on the model.
This is a special case of base-rate neglect, the cognitive bias of reasoning about case-specific information while ignoring the general frequency of the event. Applied to AI-assisted radiology, it feeds automation bias — the tendency to align excessively with the algorithm's output, documented mainly among less experienced readers. The authors cite a concrete precedent: an emergency-department tool for large vessel occlusion detection was shut down for producing "too many" false positives, despite 100% sensitivity and 92% specificity; the real culprit was the 4% prevalence.
Until now, the literature on trust calibration has worked mostly on the machine side: explainability, uncertainty quantification, better presentation of outputs. The human side — training the reader, defusing the bias before it acts — has remained marginal. Intracranial aneurysm on time-of-flight MR angiography (TOF-MRA, a sequence that makes flowing blood visible without contrast injection) is ideal terrain for testing that route: adult prevalence is around 3.2%, the lesions are small, often incidental, and routinely mimicked by normal variants — vascular loops, infundibula, perforating arteries.
The method
The setup. Prospective, multinational randomized controlled reader study, preregistered in the German Clinical Trials Register (DRKS00038740, registered 16 December 2025), approved by the ethics committee of the Technical University of Munich, patient consent waived, electronic consent from readers. Enrolment between December 2025 and May 2026.
The readers. Thirty-four radiologists from ten countries, recruited via the local department, the authors' professional networks and the European Society of Neuroradiology. Entry criterion: at least six months of brain MRI reading experience. Breakdown by level: 16 residents, 8 fellows or general radiologists, 10 certified neuroradiologists. 1:1 randomization stratified by experience level, allocated by automated script at the start of the session — 17 control, 17 intervention, with identical experience distributions in each arm. No withdrawals after randomization. The population was predominantly male (26/34), based in Germany (23/34) and academic (29/34); mean experience 8.2 ± 8.8 years (control) versus 6.7 ± 4.8 years (intervention).
The intervention. A three-minute video, a narrated slide presentation, watched before the session by the intervention arm only. It recalls the published performance of AI aneurysm detection tools, explains the false positive paradox as a prevalence-driven gap between diagnostic accuracy and predictive value, and gives realistic PPV ranges. No study case appears in it, and no case-specific training was provided. The control arm proceeded directly to the reading session with no introductory material. The video is public.
The cases. Twenty anonymized 3D TOF-MRA examinations, acquired between 2021 and 2023 at an outpatient radiology practice where a commercial aneurysm detection tool (mdbrain version 4, mediaire GmbH, Berlin) is used in routine practice. Each examination carries exactly one AI flag, mean diameter 2.2 ± 1.2 mm. The set was purposively sampled to reach a PPV of 10%: 2 true positives, 18 false positives. The false positives correspond to vascular loops (7/18), infundibula (3/18) and perforators (3/18). No AI false negatives are present. Twelve of the twenty cases come from a previous study by the same team. Readers were blinded to the case composition and received no case-level feedback.
The session. Before any reading — and, for the intervention arm, before the video — each reader estimated the PPV of an AI aneurysm detection tool on a five-point ordinal scale in 20% increments. Then, for each of the 20 cases presented in random order, they viewed the AI-annotated series and the unannotated series in sequence, then recorded (1) the likelihood of a true aneurysm on a 4-point scale from "excluded" to "certain", and (2) one of three strategies: no follow-up, MRI follow-up, or digital subtraction angiography (DSA, an invasive procedure). All readings in both arms were AI-assisted; the authors deliberately omitted an unassisted baseline to limit participant burden.
Outcomes and analysis. Two prespecified primary outcomes: acceptance rate of false positives, and follow-up intensity (ordinal scale none < MRI < DSA). Mixed models with crossed random effects for reader and case — a generalized linear mixed model for acceptance, cumulative link mixed models for ratings and follow-up, with arm and experience as fixed effects. No a priori power calculation — no comparable effect size was available — no correction for multiple comparisons, one-sided superiority tests, alpha at 0.05. The authors themselves describe the design as exploratory.
The reference PPV. To obtain a benchmark, the authors propagated by Monte Carlo simulation (2,000,000 draws) the uncertainty of three published parameters: aneurysm prevalence 3.2% [1.9–5.2], sensitivity 91.2% [82.2–95.8] and specificity 83.5% [72.9–90.6] from a meta-analysis.
The results
The opening gap. Simulation gives a PPV of 15.4% [8.1–28.0] — roughly one flag in six corresponds to a true aneurysm. The readers estimated on average 62.9% (65.3% in the control arm, 60.6% in the intervention arm). Twenty-three of 34 (67.6%) put PPV above 60%, seven (20.6%) above 80%, and only two (5.9%) below 20%. That is a fourfold error among professionals whose job this is.
First outcome: acceptance of false positives. In raw counts, across 306 false positive reads per arm (17 readers × 18 cases), the intervention group rated "aneurysm likely" 42 times (13.7%) and "certain" 19 times (6.2%), versus 62 (20.3%) and 26 (8.5%) in the control group. In the mixed model, acceptance probability came out at 12.7% [6.0–25.0] versus 22.5% [11.6–39.2], an odds ratio of 0.50 (upper 95% confidence bound: 0.95; one-sided p = 0.017). The ordinal model on ratings points the same way (OR 0.53; p = 0.036). Experience was not a significant predictor (p = 0.302), and between-case variance (σ² = 1.49) exceeded between-reader variance (σ² = 0.76): the case matters more than the person reading it.
Second outcome: follow-up intensity. For false positives, the intervention arm recommended follow-up MRI in 21.9% of cases (67/306) versus 32.0% (98/306), and DSA in 17.3% (53/306) versus 22.9% (70/306) — a follow-up recommended overall in 39.2% (120/306) versus 54.9% (168/306). OR 0.47 (upper bound 0.81; one-sided p = 0.014). Again experience was not predictive (p = 0.210), and the widest gap appeared among generalists and fellows (29 percentage points).
Secondary outcomes. Sensitivity was 97.1% (33/34) in the intervention arm versus 82.4% (28/34) in control — but those denominators are 17 readers × 2 true positives, which is to say almost nothing. Findings not flagged by the AI were judged likely or certain in 3.2% of cases (11/340) versus 4.7% (16/340). Perceived tool quality was lower in the intervention arm (mean 2.5 versus 3.2 out of 5), without reaching significance (p = 0.060); willingness to use it in routine practice did not differ (p = 0.356). Intervention readers rated the video's influence as moderate to strong (median 4/5).
Clinical translation. Take 1,000 patients whose examination carries an AI flag, at the realistic PPV of 15.4%: roughly 846 of those flags are false. With control-arm behaviour (54.9% follow-up), that would mean around 464 unnecessary follow-up examinations, including roughly 194 DSAs. With intervention-arm behaviour (39.2%), around 332 follow-ups, including roughly 146 DSAs. The gap — on the order of 130 examinations and 48 DSAs avoided per 1,000 flagged patients — is an extrapolation, not a field measurement; it gives the order of magnitude of what is at stake.
What is good
This is a genuine randomized trial, preregistered, on a free intervention. 1:1 randomization stratified by experience level, allocation by automated script, DRKS preregistration predating enrolment, prespecified primary outcomes, no withdrawals after randomization, mixed models with crossed random effects for reader and case — that is, an analysis that acknowledges the 306 reads in an arm are not 306 independent observations. In an AI-health literature where retrospective comparison on a public benchmark remains the norm, this level of methodological rigour deserves to be named.
The intervention does not touch the model. No threshold, no interface, no retraining: three minutes of generic video, transferable as-is to any rare condition and any tool. That is the paper's reusable contribution. Almost all work on trust calibration concerns what the machine displays; here the manipulated variable is what the reader knows before opening the examination — a conceptual shift that costs almost nothing to replicate. And the video itself is published.
The authors dismantle their own most marketable result. They write that the maintained sensitivity must be interpreted cautiously given the number of true positives; that the control arm may have recalibrated through repeated exposure to false positives alone — which would attenuate the measured effect relative to real practice; that the absence of an unassisted baseline prevents quantifying automation bias in absolute terms; and that the study is underpowered to detect undertrust induced by the video. They also flag the ironic effect: the same video that improves calibration lowers the perceived quality of the tool. No funding, and the software manufacturer played no role in design, case selection, analysis or writing, and did not review the manuscript.
What is less good
The case set makes the primary outcome nearly free — and its cost invisible. Ninety percent of the flags presented are false. In that regime, anything that makes a reader more sceptical mechanically improves the primary outcome: a threshold shift suffices, with no improvement in discrimination (the ability to separate true from false). The price of that scepticism would be paid on true positives, and there are two. The 97.1% versus 82.4% sensitivity rests on 34 reads per arm and a difference of one missed aneurysm. A design that measures benefit over 306 observations and cost over 34 cannot conclude on net balance. This is the misleading metric failure mode, applied not to a model but to a protocol.
The comparator is "nothing at all", not another intervention. The control arm received no introductory material. The measured effect therefore mixes the video's content with everything that comes with it: priming, attention to the task, and above all a demand effect that is hard to rule out — a reader who has just watched three minutes on false alarms and immediately moves on to a reading session can guess what is expected. The authors note that participants were unaware of the hypothesis and of the other arm's existence, but the intervention arm was obviously not blinded to its own allocation. An active comparator — a video of equal length on normal variants, or PPV displayed in the interface — would have separated the mechanism from its packaging. As with AI-embedded abdominal ultrasound, the choice of comparator determines half the result here.
Modest sample, narrow population, closed data. Thirty-four readers, 23 of them in Germany (67.6%) and 29 in academic institutions (85.3%): generalization to community radiology, other training systems or other cultures of AI use is neither tested nor established — a classic population bias, here on the human side of the evaluation. No a priori power calculation, one-sided tests, no multiplicity correction, and upper confidence bounds of 0.95 and 0.81 that hug unity: both primary outcomes are statistically significant, but fragile. Twelve of the twenty cases come from a previous study by the same team. Finally, reader-level data and statistical code are available only "on reasonable request to the authors" — that is, not open. One author declares being a shareholder in Bonescreen GmbH and receiving speaker honoraria from Bracco and Novartis; the others declare none.
What it changes
For the research community. This paper moves the intervention variable. As long as one works on what the model displays — saliency map, confidence score, uncertainty interval — one assumes the problem is an information deficit. Here the hypothesis tested is that the problem is a framing deficit: readers already had sensitivity and specificity, and were still off by a factor of four on PPV, because those two numbers are not enough to compute it. The natural follow-up is a design this study does not provide: an active comparator, a case set at realistic rather than enriched prevalence, enough true positives to measure the cost, and a delayed measurement — how long do three minutes of video last? The paper does not say, and that is the most important question for anyone considering turning this into training.
For clinicians. The immediately usable result is not the video, it is the opening number: 62.9% versus 15.4%. It means the average radiologist overestimates by a factor of four the reliability of an AI flag on a rare condition, and there is no reason to think this stops at aneurysms — a recent cross-sectional analysis of 38 FDA-authorized radiology AI devices finds clinically anticipated PPVs below 50% for the majority of target pathologies, despite sensitivity and specificity both above 90%. The authors' two concrete recommendations are to display a predictive value range alongside the model's output — a public calculator from the American College of Radiology exists for this — and to trigger the AI on cohorts with higher pretest probability rather than on every technically eligible examination. The logic is the same as for calibrating a prognostic score in intensive care: discrimination performance says nothing about what to do until it has been crossed with prevalence.
For patients and the public. The dominant intuition holds that a medical AI "90% accurate" is wrong one time in ten. This paper shows that on a rare disease, the same AI can be wrong five times out of six when it flags something — and that professionals themselves do not know it. The consequence is not a missed diagnosis, it is the opposite: follow-ups, repeat MRIs, angiograms, anxiety, for perfectly normal vascular loops. That is the administrative and anxiety-producing face of AI in health, less spectacular than a diagnostic error but far more frequent. The same mechanics explain why headline performance in mammography screening degrades as soon as one leaves benchmark conditions. And there remains the trade-off the authors own: better informing radiologists makes them less enthusiastic about tools that, in a higher-prevalence setting, would be useful. Calibrated trust is not always greater trust.
Further reading
The preprint: A 3-Minute Education on the False Positive Paradox Improves Trust Calibration in AI-Assisted Intracranial Aneurysm Detection: A Multinational Randomized Controlled Reader Study, medRxiv, DOI 10.64898/2026.08.25.26361324, posted 28 August 2026. Authors: Su Hwan Kim, Bastien Le Guellec, Paula Roßmüller, Severin Schramm et al., with Benedikt Wiestler and Dennis M. Hedderich as senior authors — Technical University of Munich, Lille University Hospital, RWTH Aachen, Cologne, Regensburg, Berlin, Brown University. Preregistered study: DRKS00038740. The three-minute video used as the intervention is publicly available. The software evaluated is mdbrain version 4 (mediaire GmbH, Berlin); the manufacturer played no role in the study. No funding. Reader-level data and statistical code are available on request from the corresponding author, but are not published.