Predicting preterm birth from a cervical ultrasound: why these AI models collapse on a different scanner

A team from the GARBH-Ini cohort (North India), with the University of Oxford, trained several models — including a Vision Transformer — to predict spontaneous preterm birth from a single mid-trimester cervical ultrasound. Internally, the best model reached an AUC of 0.71; tested on an independent cohort acquired with a different scanner, it fell to 0.52, barely above chance. A negative result, honestly published, that mostly says what medical imaging AI cannot yet do: walk out of the lab.

The context

Prematurity — birth before 37 weeks of gestation — remains the leading cause of neonatal mortality worldwide, and some of these deliveries occur without an identifiable trigger: this is called spontaneous preterm birth. Flagging at-risk pregnancies in advance would allow action (vaginal progesterone, cerclage, closer monitoring). Today's reference tool is measuring cervical length by ultrasound in the second trimester: a short cervix signals raised risk. But that marker has modest sensitivity — it misses many women who will nonetheless deliver early.

Hence the hypothesis behind this work: the image of the cervix may hold more information than its length alone. The texture of the tissue — its grain, its micro-variations of echogenicity — might carry a signature of cervical remodelling invisible to the eye and not summarised by a distance. That is exactly the kind of pattern a vision model is meant to extract. One question decides everything in medical imaging: does a model that works on one centre's images still work on another device's, in another hospital? A reading cue: the AUC (area under the ROC curve) measures the ability to tell a positive case from a negative one; 1.0 is perfect, 0.5 equals a coin toss.

The method

The data come from GARBH-Ini, a prospective pregnancy cohort followed in North India — hence a population under-represented in the imaging literature, which matters. Each participant had a second-trimester cervical ultrasound, and the outcome to predict is spontaneous delivery before 37 weeks. The team tests not one approach but four families, allowing classical and deep methods to be compared on the same ground.

The first is a hand-crafted texture descriptor: Local Binary Patterns (LBP), which for each pixel encode the neighbourhood pattern — brighter or darker around it — summarising the image into a histogram of micro-textures, then classified by a random forest (an ensemble of decision trees voting by majority). The second is a Vision Transformer (ViT), a neural network that splits the image into tiles and learns, through an attention mechanism, which regions matter relative to one another — the architecture now dominant in vision. The third rests on clinical variables alone. The fourth is multimodal: it fuses image and clinical data.

The crucial point is the evaluation protocol. A model is first judged on an internal test set (patients from the same cohort, held out). Then comes the test that truly counts, external validation: the model faces an independent cohort, acquired with a different scanner. This is the test that simulates the reality of deployment — a tool sold to a hospital lands on machines, operators and populations it never saw during training.

The results

Internally, the signal exists. The best model reaches an AUC of 0.71 (95% confidence interval: 0.60–0.82), slightly above what cervical length alone gives. In other words, on its own cohort's data the model captures something the classical measure does not fully summarise. This is the kind of figure that, in a press release, would become "an AI predicts preterm birth".

External validation undoes that reading. On the independent cohort acquired with another device, performance collapses to an AUC of 0.52 (95% CI: 0.38–0.64) — statistically indistinguishable from chance. Neither the Vision Transformer nor the multimodal model does better: sophistication did not protect against the fall. A clinically high-risk subgroup, at the 34-week threshold, shows somewhat better discrimination, but on too few cases to conclude. The clinical translation is direct and severe: an AUC of 0.52 means that at the patient's bedside, on a new machine, the model cannot separate those who will deliver early from the rest. Using it as is would either falsely reassure women genuinely at risk, or trigger interventions (cerclage, progesterone, admission) on pregnancies that do not warrant them.

What is good

An external validation done, and published despite the failure. The vast majority of imaging studies stop at the internal test and headline the flattering AUC. Here the team went as far as the independent cohort on another scanner, watched its model collapse, and reported it anyway. This negative result is worth more, for the field, than ten unreplicated successes: it pinpoints exactly where the wall stands.

An honest comparison, several models against a real benchmark. Testing in parallel a classical texture method (LBP), a Vision Transformer, a clinical model and a multimodal fusion, and pitting them against cervical length — the standard of care — avoids the usual "our model versus nothing" bias. It also shows the failure is not that of a particular architecture but of the problem as posed.

A prospective cohort and an under-represented population. GARBH-Ini is a prospective pregnancy follow-up in North India, where imaging AI research is overwhelmingly trained on North American and European data. Asking the prematurity question in this population, with a dedicated collection rather than an opportunistic set, is a sound methodological choice.

What is less good

Very likely shortcut learning. The sharp drop from 0.71 internally to 0.52 externally, the moment the scanner changes, is the signature of shortcut learning: the model most likely learned the acquisition signature of the training machine — settings, proprietary post-processing, characteristic grain — rather than cervical biology. It was not recognising an at-risk cervix, it was recognising a device. This is the central failure mode of imaging AI, and it is demonstrated here at full scale.

A biologically heterogeneous target, and a fragile subgroup metric. Spontaneous preterm birth is not a single disease but a syndrome with multiple causes (infection, inflammation, mechanical factors, idiopathic). Expecting a single cervical image to predict all these trajectories is probably unrealistic. And the best apparent result — the high-risk subgroup at the 34-week threshold — rests on few cases: it is a misleading metric, an encouraging figure too imprecise to rely on.

Evidence limits to keep in mind. This is a not-yet-peer-reviewed preprint; the external confidence intervals are wide (0.38–0.64), a sign of limited validation numbers; only one other machine was tested, whereas true generalisation would require several manufacturers and centres; and no code or weights release is announced, which for now prevents replaying the audit elsewhere. One therefore cannot rule out that a more robust training protocol would do better — the study shows the failure of one approach, not the impossibility of the task.

What it changes

For the research community, the message is a salutary reminder: an internal AUC, however honest, is worthless until it has survived a change of device. Multi-machine external validation should be the norm, not the exception, and negative results like this one deserve to be published far more often. The concrete work now opening is hardware robustness: image harmonisation, data augmentation simulating several scanners, domain-invariant learning.

For clinicians, nothing changes in practice, and that is the useful information: none of these models is ready for the bedside. Cervical length measured by ultrasound, combined where appropriate with vaginal progesterone, remains the reference framework to screen for and prevent prematurity. A cervical-texture AI is not, today, a reliable complement.

For patients and the public, the lesson is an antidote to marketing. When an announcement promises an "AI that predicts preterm birth", the right question is not "how accurate is it?" but "was it tested on another device, in another hospital, on another population?". This study shows how far the answer can move a tool from promising to useless — and why the caution of teams that refuse to deploy too fast here directly protects pregnant women.

To go further

The preprint is available on medRxiv (2026.07.17.26358221), posted in July 2026 by Radhika Chanian, Divyanshu Mishra, Rahul Jain, Nikhil Sharma, Ashok Khurana, Reva Tripathi, Abhinav Tripathi, the GARBH-Ini study group, Nitya Wadhwa, J. Alison Noble (University of Oxford), Ramachandran Thiruvengadam, Bapu Koundinya Desiraju and Shinjini Bhatnagar (THSTI); it has not yet been peer-reviewed. On the generalisation and external validation of imaging models, see our decryptages on the external clinical validation of a skin-cancer dermoscopy AI, on the robustness of pathology foundation models to perturbations, and on the multi-site validation of a stroke-risk model.

Editorial transparency: French version written and signed by the Tatakoto editorial team based on a reading of the preprint. English, Spanish and Chinese translations produced with AI assistance and reviewed.