Tuberculosis on chest X-rays: medical vision-language models collapse once healthy controls are replaced by sick ones

Three researchers at the Indian Institute of Technology Indore put four vision-language models — three medical (BioMedCLIP, CheXficient, MedSigLIP) and one general-purpose (OpenCLIP) — through a tuberculosis-screening audit on 12,200 chest radiographs from four public datasets, changing one thing at a time: the cohort, the wording of the text, the kind of controls, the assumed prevalence, the decision threshold. The rankings hold up poorly: replacing healthy controls with patients who have other diseases cuts AUROC by 0.075 to 0.306, and against pneumonia or lung tumour the three medical models fall below 0.5, the level of chance; a supervised network at 0.999 on internal validation drops to 0.629 on two external cohorts. The paper proposes no new model, and that is its strength: it shows that a benchmark score is not a property of a model, but of a complete evaluation specification.

The context

Tuberculosis remains one of the leading infectious causes of death worldwide, and the chest X-ray is the cheapest mass-screening tool available. Since 2021, the World Health Organization has accepted computer-aided detection (CAD) as an alternative to human reading for screening and triage in people aged 15 and over. Several commercial products are already deployed in high-burden countries where radiologists are scarce.

Those products are supervised classifiers: trained on thousands of images labelled "tuberculosis" or "not tuberculosis". Vision-language models (VLMs) promise something else. An image encoder and a text encoder are trained together on image-caption pairs so that a radiograph and a sentence land in the same mathematical space. An image can then be classified zero-shot, meaning without any retraining: you write "a chest X-ray showing tuberculosis" and "a chest X-ray without tuberculosis", and check which sentence the image sits closer to. For a screening programme with no local labelled data, the argument is appealing.

The question is what a benchmark actually measures. A high AUROC says nothing about which patients served as controls, which sentence defined the classifier, or whether a threshold set elsewhere will work here. The authors describe any evaluation by seven components — model, text, cohort, definition of negatives, prevalence, threshold, metric — and changing one means evaluating a different system.

The method

The data. Four public datasets. Montgomery (United States, 58 tuberculosis, 80 normal) and Shenzhen (China, 336 tuberculosis, 326 normal), the two long-standing National Library of Medicine collections. TBX11K, which has the rare merit of distinguishing three categories: tuberculosis, healthy, and sick but not tuberculous; its validation split holds 200 tuberculosis, 800 healthy and 800 sick images. Finally the VinDr-CXR test set (3,000 images from two Vietnamese hospitals, labelled by consensus of five radiologists), which names the competing diagnoses: 164 tuberculosis, 2,051 no-finding radiographs, 210 pneumonias and 73 lung tumours without tuberculosis.

The models, all frozen. OpenCLIP, trained on two billion web images, is the general-purpose comparator. BioMedCLIP (Microsoft) learned from 15 million figure-caption pairs from the biomedical literature. CheXficient (Stanford AIMI) specialises in chest radiography — and VinDr-CXR is explicitly among its training sources, which the authors flag throughout. MedSigLIP (Google) covers several medical specialties.

The prompts. Five families of wording, published verbatim: a "clinical" family of four sentences per class, designated in advance as primary, and four single-sentence variants (report style, presence/absence, and so on). In total, 244,000 scores.

The analyses, fixed before looking at the results. Discrimination via AUROC — the probability that a tuberculosis image receives a higher score than a non-tuberculosis image; 0.5 is chance. Score reliability via the Brier score (mean squared error between score and outcome) and calibration error (the gap between stated probability and observed frequency). Then two operational tests: reweighting the classes to simulate different prevalences, and transporting a threshold chosen to reach 95% sensitivity on TBX11K to the other cohorts, untouched. Finally, a conventional supervised ResNet-50, trained on TBX11K with five random seeds, serves as reference; the authors also screen for near-duplicate images across datasets using perceptual hashing.

The results

No model wins everywhere. With the clinical prompt, MedSigLIP edges CheXficient on Montgomery (0.976 vs 0.969) and TBX11K (0.800 vs 0.795), without a statistically resolved difference; CheXficient clearly wins on Shenzhen (0.950 vs 0.907). On VinDr-CXR, CheXficient reaches 0.878, MedSigLIP 0.829, BioMedCLIP 0.713, OpenCLIP 0.528 — but CheXficient saw that dataset during training. On TBX11K, AUROCs are nearly identical while CheXficient's Brier score is almost twice as good (0.099 vs 0.171).

The text is part of the model. Switching prompt family significantly changes AUROC in 21 of 48 comparisons after correction for multiple testing. CheXficient goes from 0.835 with "report" wording to 0.622 with "presence/absence" on the same TBX11K images.

The central result: the controls. Holding the 200 tuberculosis images fixed, replacing 800 healthy controls with 800 sick controls lowers AUROC for every model: MedSigLIP from 0.953 to 0.647, BioMedCLIP from 0.872 to 0.611, CheXficient from 0.865 to 0.725. On VinDr-CXR the effect is starker. Against no-finding radiographs, CheXficient reaches 0.945; against pneumonia, 0.466; against lung tumour, 0.372. MedSigLIP and BioMedCLIP fall to between 0.31 and 0.38. Put differently, when facing another disease, these models more often give a high "tuberculosis" score to the pneumonia image than to the tuberculosis image.

Thresholds do not travel. A threshold tuned for 95% sensitivity on TBX11K keeps that sensitivity in only 4 of 16 transports. And when it does, it is sometimes by calling almost everyone positive: OpenCLIP keeps its sensitivity on Shenzhen with a specificity of 0.043. CheXficient on VinDr-CXR reaches 97.6% sensitivity and 35.4% specificity; against pneumonia it correctly rejects only one of 210, and none of the 73 tumours.

Clinical translation (Tatakoto's calculation from these figures, not an estimate by the authors): out of 1,000 radiographs at the prevalence observed in VinDr-CXR (5.5%), this setting would catch about 54 of 55 tuberculosis cases, but would also send about 610 people without tuberculosis for confirmatory testing. For comparison, the WHO target product profile for a triage test calls for at least 90% sensitivity and 70% specificity.

Supervision does not save the day. The ResNet-50 reaches 0.997 on internal validation (0.999 for the five-seed ensemble), then 0.629 on both Shenzhen and Montgomery. Perceptual-hash screening flags 249 suspicious pairs involving 313 images; removing them and retraining improves Shenzhen from 0.633 to 0.733 and Montgomery from 0.627 to 0.678, without closing gaps of about 0.27 and 0.32.

What is good

Paired comparisons that isolate a single factor. The healthy-versus-sick control test keeps exactly the same 200 tuberculosis images and the same prevalence (20%); only the nature of the negatives changes. That is what allows the AUROC drop to be attributed to the control spectrum rather than to a change in the positive population. All four differences are significant after Holm correction (p ≈ 0.0004).

Rare transparency about what could skew the audit. Prompts are published verbatim, exact weight revisions are given for two of the four models, CheXficient's exposure to VinDr-CXR is flagged in every table, and the authors state that their MedSigLIP preprocessing differs from the developer's (PIL resizing rather than TensorFlow), so their numbers do not reproduce the model card.

Separating three claims the literature conflates. "This model ranks well", "its scores are reliable" and "its threshold keeps its sensitivity elsewhere" are three distinct propositions, and the audit shows they can diverge for the same model. Example: at a specified prevalence of 50%, MedSigLIP has a better Brier score than CheXficient on all three cohorts; at 10%, the order reverses — without a single score changing.

What is less good

Labels that are not microbiological truth. None of the four datasets provides harmonised bacteriological confirmation; VinDr-CXR rests on radiological impressions. Some of the "failures" against pneumonia may reflect genuinely ambiguous images or debatable labelling rather than model error. The authors acknowledge it: AUROCs below 0.5 describe a ranking, not a clinical misdiagnosis rate. This is the classic spectrum bias problem applied to the reference standard itself.

The unit is the image, not the patient. The manifests do not link images from the same patient, and the duplicate screen excludes VinDr-CXR and cannot certify that no evaluation image was used to pre-train these models, whose corpora are partly non-public. The risk of data leakage is bounded, not ruled out, and confidence intervals are probably too narrow. The supervised 0.999 → 0.629 gap in particular leaves room for interpretation: the authors show that removing suspected near-duplicates helps somewhat, without being able to say which were true duplicates.

An audit that stops short of the clinical question. No reader study, no fairness analysis by sex or age, no evaluation of a locally recalibrated system — which is what any serious screening programme would do. The comparators are generic VLMs and a ResNet-50, not the commercial CAD products actually deployed for tuberculosis: the audit says nothing about how they behave against the same sick controls. Finally, no statement on funding, conflicts of interest or released code appears in the version reviewed, which prevents rerunning the analysis as is.

What it changes

For the research community, the paper supplies a directly reusable checklist: any zero-shot medical VLM publication should give its exact prompts, its definition of negatives, and report performance against named sick controls, not only healthy ones. It is a tooled-up version of a known problem, hidden stratification — an overall AUROC of 0.878 on VinDr-CXR can coexist with worse-than-chance ranking against pneumonia, because 2,051 normal images dominate the denominator. The same mechanism appeared in our analysis of paediatric chest X-ray model transportability and in the one on auditing MIMIC-CXR labels.

For clinicians and screening programmes, the message is practical. In a real population, screened people are not "tuberculous or healthy": many have pneumonia, sequelae, cancer. A model presented with a flattering AUROC on Montgomery or Shenzhen, which compare tuberculosis only with normal lungs, provides no evidence for that scenario. And a threshold should never be imported as is: it must be reset and validated locally, with sensitivity and specificity reported together.

For patients and the public, a tool that confuses pneumonia with tuberculosis overloads laboratories with confirmatory tests and delays other diagnoses. The zero-shot promise — a model usable anywhere without local data — is not demonstrated here for tuberculosis.

Further reading

The paper: Beyond Benchmark Scores: Auditing Medical Vision-Language Models for Chest X-Ray Tuberculosis Screening, Mushir Akhtar, M. Tanveer, Mohd. Arshad (Department of Mathematics, Indian Institute of Technology Indore), arXiv:2609.21763 (cs.CV), submitted 18 September 2026, DOI 10.48550/arXiv.2609.21763. Preprint, not peer reviewed.

The datasets: Montgomery and Shenzhen (National Library of Medicine), TBX11K, VinDr-CXR v1.0.0 (PhysioNet). On the WHO framework: the consolidated guidelines on tuberculosis screening (2021). On another case of spurious correlation tied to data origin: shortcut learning in mammography.