Extubation failure: what respiratory therapy notes read by an LLM actually add — 1.9 AUROC points, and nothing at all the following year
A University of Washington team has Llama-3-8B read the respiratory therapy notes from the twelve hours before an extubation, extracts nine interpretable clinical variables from them — sputum quantity and thickness, cough strength, cuff leak — and adds them to a tabular model predicting extubation failure across 3,243 intensive care stays. The gain is 1.9 AUROC points, from 0.733 to 0.752, with confidence intervals that overlap almost entirely and no significance test on the difference. On the paper's only temporal split — train on 2021-2022, test on 2023 — the gain reverses, and that line, relegated to an appendix, is the one worth reading first.
The context
Extubating an intensive care patient means removing the breathing tube and handing them back their own respiration. The decision is made daily, and it is costly in both directions. Too early: the patient deteriorates, reintubation is required, and reintubation is associated with higher mortality and longer stays. Too late: every additional day of invasive ventilation adds risk of ventilator-associated pneumonia, prolonged sedation and ICU-acquired weakness.
The classical weaning toolkit is well known: the spontaneous breathing trial, the rapid shallow breathing index, the cuff leak test. Machine learning models published over the past decade report AUROCs of 0.83 to 0.85 — but on cohorts whose inclusion criteria differ enough to make the comparison meaningless, which the authors take the trouble to tabulate. Some require no minimum ventilation duration, others twenty-four hours; some define failure at two days, others at seven.
The gap this paper targets lies elsewhere. Part of the information clinicians actually rely on at the bedside exists in no structured field of the record: the volume of secretions, their thickness, the vigour of the cough, the presence of a leak. That information lives only in the free text of respiratory therapy notes, written several times a day by respiratory therapists, a profession specific to the North American health system. The question posed is therefore precise: if that text is turned into variables, does the model improve?
The method
The cohort. An electronic health record assembled across the three hospitals of University of Washington Medicine, from April 2021 to September 2023. Initially, 10,194 ventilated patients over 10,810 encounters, and 34,834 respiratory therapy notes preceding first extubation. After exclusions — 843 for missing height or weight, 4,711 with no invasive ventilation episode of at least twenty-four hours, 1,922 carrying a do-not-intubate or do-not-resuscitate order within seven days of extubation, 91 duplicate patients — 3,243 encounters remain. Extubation failure is defined as death or return to invasive ventilation within seven days: 647 cases, or 19.95% — 618 reintubations and 29 deaths.
The extraction. The model used is Meta-Llama-3-8B-Instruct, open-weight, at temperature 0.01. An important point: this is neither embeddings — those opaque numerical representations of a text — nor named entity recognition, but whole-note classification by few-shot prompting. A separate prompt per concept, a constrained answer (yes/no, low/medium/high), and an "unlabeled" value when the concept is not mentioned. Fifteen variables are defined from the literature and from pulmonologists' input, across five families: suctioning, sputum, cough, breath sounds, cuff leak.
The quality of that extraction is measured before any use, which is not the norm: 400 hand-annotated notes, 200 to build the prompts and 200 to evaluate them, reviewed and corrected by two practising critical care pulmonologists. Macro-F1 per variable ranges from 0.53 for spontaneous cough to 0.99 for thick sputum. Six variables are discarded for F1 below 0.8 or for collinearity — suctioning presence correlates at 0.808 with sputum presence. Nine survive.
The windows. The notes retained are those from the twelve hours before extubation: 2,869 notes, covering 2,509 of the 3,243 encounters — a quarter of the cohort therefore has no usable note. Structured variables are averaged over the preceding four hours: vital signs, laboratory values, ventilation parameters, medications, thirty-two admission diagnoses, age over 60, documented male sex.
The models. Logistic regression and gradient boosting, plus two survival models (Cox and gradient boosting survival) binarised at seven days. Split by stratified random sampling on documented race or ethnicity and on outcome: 646 test encounters, or 20%. Eight-fold cross-validation on the training split only, solely for hyperparameter tuning. Decision threshold fixed at the training prevalence, 0.201.
The results
The raw figures. On the test set, logistic regression on structured variables alone reaches an AUROC of 0.729 (95% CI 0.684-0.775); with medications, 0.733 (0.684-0.783); with notes, 0.749 (0.701-0.795); with both, 0.752 (0.703-0.794). AUPRC — the area under the precision-recall curve, more informative than AUROC when positives are rare — moves from 0.388 to 0.399. Gradient boosting is uniformly worse, between 0.669 and 0.671, on account of manifest overfitting (0.782 on train against 0.712 for logistic regression). Notes alone give 0.605 (0.548-0.658); ventilation duration alone gives 0.658.
The incumbent comparator. The hospital already deploys a decision support checklist. On the test set it achieves 0.55 sensitivity, 0.55 specificity and 21% precision. The model raises precision to 34%. Of the 81 failure patients actually assessed by the checklist, it flags 40, the model flags 59, and 32 are common to both — the authors conclude, correctly, that the two tools are complementary rather than substitutable.
The clinical translation. At the authors' own threshold, out of 1,000 extubations of which 200 fail: the model without notes detects roughly 139 failures at the cost of 264 false alarms; with notes, roughly 147 failures for 268 false alarms. Eight more patients caught, four more alarms. At matched sensitivity, however — 74.2%, meaning 5% of failures missed — the false positive rate falls from 40.0% to 34.1%, which the authors translate as "five fewer false alarms per 100 patients". Decision curve analysis gives a net benefit of about 0.01, that is one additional true positive per hundred patients treated, between thresholds 0.17 and 0.22.
What the model sees. The logistic regression coefficients, all variables normalised, put initial ventilation duration first (0.0955), then plateau pressure (0.0679), SpO2 (−0.0664) and urea nitrogen (0.0601). Sputum quantity comes thirteenth (0.0392), ahead of age and FiO2. It is the only note-derived variable to enter the top fifteen coefficients, alongside sputum thickness and weak cough in the broader model.
What is good
The extraction is validated before it is used, and six variables out of fifteen are thrown out. This is the methodological move that sets this paper apart from the mass of work that wires an LLM output straight into a classifier. Here: 400 hand-annotated notes, reviewed by two intensivists, a published F1 for each of the fifteen variables, and an elimination rule applied even when it costs — suctioning presence, at F1 0.873, is dropped for collinearity with sputum. The choice of classified variables over dense embeddings is owned as an explicit trade of expressiveness for interpretability, which is the honest position. The same trade-off arose in extracting a frailty index from clinical notes.
Every analysis usually skipped is performed. Calibration reported — Brier score of roughly 0.14, curve slope above 1 for both variants, and the authors themselves note that the note features do not improve it. Decision curve analysis, which quantifies net benefit rather than assuming it. A subgroup performance table. An ablation shuffling the note columns, which does make the gain vanish. And above all two sensitivity analyses missing almost everywhere else: on the inclusion threshold (one hour versus twenty-four hours of minimum ventilation) and on the event definition window, from twelve hours to fourteen days.
The comparator is the tool actually deployed, and the comparison is reported without arrangement. Many papers measure themselves against an obsolete score or against nothing. This one measures itself against the checklist the hospital uses, publishes the patient-by-patient overlap table, and draws the conclusion unfavourable to the model: 32 of the 40 patients flagged by the checklist are also flagged by the model, so one does not replace the other. The code is released, final prompts included.
What is less good
The gain sits inside the noise, and Appendix H shows it. This is the central point. The 1.9-point difference separates two confidence intervals that overlap along almost their entire length: 0.684-0.783 against 0.703-0.794. No significance test is reported on that difference — no DeLong test, no paired bootstrap — even though Table 4 lines up fourteen arms, which mechanically multiplies the chances that some gap appears. And the paper supplies its own counterexample: on the temporal split, training on 2021-2022 and testing on 2023, AUROC is 0.751 with the extracted variables and 0.756 without. The gain reverses. On a random subsample of the same size it returns: 0.761 against 0.754. In other words, the entire effect fits inside the difference between two ways of cutting the same dataset. This is the misleading metric in its most ordinary form: an exact figure, an honest interval, and a conclusion the interval does not carry.
No classical text baseline, and no outside centre. The gain is established against structured data alone. Neither TF-IDF, nor bag-of-words, nor dense note embeddings are tested — yet that is exactly the question one asks on reading the title: are these variables better than a method from 1995? The authors mention embeddings in the discussion, to say they would be less interpretable, never to evaluate them. As for the clinical comparator, a checklist at 0.55 sensitivity and 0.55 specificity is barely better than a coin toss: beating it establishes little. And there is no external or prospective validation — the authors justify this honestly ("no comparable dataset is publicly available", public intensive care databases contain no respiratory therapy notes), but the justification changes nothing about the fact: transportability is entirely untested.
Population bias, and a subgroup gap that cannot be settled. One health system, three hospitals in a single American city, English-only notes — the authors themselves flag that the pipeline might not hold in another language. The subgroup table gives an AUROC of 0.673 among White patients (n = 2,009, by far the largest group) against 0.820 among Asian patients (n = 225). The model works worst on the group it has the most data for, which is not the usual direction of bias and is not explained. The authors state that no difference is significant across 1,000 resamples and that the test set is too small to detect one: that is an admission the fairness question remains open, not evidence of fairness. Add the exclusion of the 1,922 encounters carrying a do-not-resuscitate order — nearly one encounter in five, and precisely those where the death component of the outcome would have fired — and the acknowledged absence of treatment counterfactuals: high-risk patients were treated, and nothing in this design separates risk from the response to risk.
What it changes
For the research community, two lessons, one of which outgrows the subject. First: the "LLM as structured-feature extractor" pattern pays off mainly in interpretability, not accuracy. The signal specific to the notes exists — 0.605 AUROC on their own, and sputum quantity entering the top fifteen coefficients — but it is largely redundant with the tabular data. Second, and more important: this paper contains, in an appendix, the demonstration that its own headline result does not survive a temporal split. Any group announcing two AUROC points without a temporal split or an outside centre is announcing noise until proven otherwise, and that Appendix H is the template for the analysis peer review ought to demand.
For clinicians, nothing changes at the bedside. Two figures are nonetheless worth keeping by anyone running a weaning alert. First, alert burden: at matched sensitivity, about five fewer false alarms per hundred extubations — modest, but the right unit, because alert burden is what kills deployed intensive care models long before AUROC does. Second, the effect of inclusion criteria: the same model reads 0.796 when short intubations are included and 0.752 when they are not, with the failure rate moving from 11.3% to 20.0%. No published extubation AUROC is interpretable without its inclusion criteria, and the stratification shows it starkly: the one-hour-threshold model recalls 85.5% of failures among patients ventilated more than twenty-four hours, and 25.8% among the rest.
For patients and the public, a simple idea. Removing the breathing tube is a judgement call, made daily, that goes wrong roughly one time in five in this cohort. An AI reading the therapist's notes does not decide: at best it flags. And the honest state of the art is that it flags about eight more patients per thousand than the record's numerical data alone, at the cost of four additional false alarms, in one hospital system, in one language, without any other hospital ever having tried it. That is not nothing, and it is not what a press release would make of it.
Further reading
The paper: Enhancing Extubation Failure Prediction with LLM-Derived Features from Respiratory Therapy Clinical Notes, Izzy Chaiken, Aditya Khowal, Neha A. Sathe, Mark M. Wurfel and Lucy Lu Wang, posted to arXiv in September 2026, categories cs.CL, cs.LG and physics.soc-ph, DOI 10.48550/arXiv.2609.17532. Published at the Conference on Health, Inference, and Learning (CHIL) 2026, PMLR volume 333. Preprint under CC BY 4.0.
Affiliations: Information School, Paul G. Allen School of Computer Science and Engineering, and Department of Medicine, University of Washington, Seattle. The third and fourth authors are practising critical care pulmonologists at UW Medicine and contributed to the classification scheme and interpretation. Protocol approved by the University of Washington Human Subjects Division under reference STUDY00018582.
Code and data: training and evaluation code, together with the final prompts, are at github.com/larchlab/extubation-failure-camera-ready. The code licence is not stated in the paper. The data are not released: they contain identifying information and reside on a server compliant with US health data privacy regulation. No weights are released, the model not having been fine-tuned — Llama-3-8B-Instruct is used as is, by few-shot prompting.
Funding: University of Washington Institute of Medical Data Science Pilot Award, the University of Washington eScience Institute cloud credits programme, gift funds from the Allen Institute for AI, and an NSF CSGrad4US Fellowship for the first author. The paper carries no conflict-of-interest declaration.
Also on Tatakoto: the transportability of an intensive care model from one database to another, decision curve analysis applied to intensive care mortality, and a prospective validation of frozen models in intensive care.