Predicting the day's anatomy before scanning: a "digital twin" for adaptive head-and-neck proton therapy

Yizhou Wu, Yuheng Li, Xiaofeng Yang, and Chih-Wei Chang propose a "digital twin" that predicts a head-and-neck patient's anatomy on treatment day, before acquiring any image, by transferring onto that patient's planning CT the way other patients changed during radiotherapy. On 88 patients, these predicted CTs match the real day's anatomy better than the frozen planning CT: +22.8% image correlation, +20.2% organ-at-risk overlap, and −23.4% CT-number error. Anticipating anatomy to speed up adaptive proton therapy is a clever idea, but the demonstration is limited to image-similarity metrics and never evaluates delivered dose — yet it is dose, not resemblance, that decides benefit in proton therapy.

The context

Proton therapy is a form of radiotherapy that uses protons instead of X-rays. Its appeal rests on a physical phenomenon called the Bragg peak: a proton beam deposits most of its energy at a precise depth, then stops abruptly. In principle, one can concentrate dose on the tumor and spare the tissue right behind it. But that precision cuts both ways: if the anatomy changes, the Bragg peak ends up in the wrong place — underdosing the tumor or overdosing a neighboring organ.

Head-and-neck cancers are the textbook case of this problem. Treatment spans four to six weeks, during which the tumor shrinks, the patient loses weight, and positioning varies day to day. A few millimeters of shift are enough to push dose toward highly sensitive organs at risk: the parotids (salivary glands, whose irradiation causes lasting dry mouth), the oral cavity, the brainstem, and the spinal cord.

The clinical answer is adaptive radiotherapy: re-planning on the day's anatomy. In practice this is most often "offline" — a new CT must be acquired, then about a week of preparation before the corrected plan is ready. The ideal "online" version adapts the plan the same day, but it runs into the time and cost of daily imaging and re-planning. This work slips into that blind spot: what if the day's anatomy could be predicted before the patient is even scanned?

The method

The basic tool is deformable image registration: a technique that warps one CT to match another and, along the way, produces a "deformation field" describing how each point moved. The authors rely on a foundation-model registration network — a model pretrained on large amounts of imaging, reused here as-is, without patient-specific retraining (a "zero-shot" approach).

The mechanism has two steps. First, take a prior patient (already treated, whose evolution is known) and register their planning CT to the target patient's, to align the two anatomies. Then estimate how that prior patient changed between planning and a scan acquired during treatment — a QACT (quality-assurance CT, taken during the sessions). This "change trajectory" is then applied to the target patient's own planning CT, producing a predicted CT (pdCT) with its organ contours automatically propagated. In plain terms: you borrow a past patient's way of having evolved and stamp it onto the new patient's starting anatomy.

To evaluate the method, the authors have 88 head-and-neck patients, each with a planning CT and three QACTs spread across treatment. The comparator is the frozen planning CT — the starting anatomy, as it would be used if nothing were adapted. Three metrics are reported. Normalized cross-correlation (NCC) measures the overall similarity between two images. The Dice score measures the overlap between two organ contours (1.0 = perfect overlap). The CT-number error quantifies the gap in density values, expressed in Hounsfield units — crucial in proton therapy, because these values are used to compute where the protons will stop.

The results

Compared with the frozen planning CT, the predicted CTs move clearly closer to the real day's anatomy: image correlation improves by 22.8%, the organ-at-risk Dice score by 20.2%, and the CT-number error drops by 23.4%. The authors note that these gains are largest for patients who changed a lot, and negligible when anatomy stayed stable — which is consistent: when nothing moves, predicting a change adds nothing.

These figures must be read for what they are: geometric measures of image resemblance and contour overlap. But in proton therapy it is not resemblance that treats a patient, it is dose. The honest clinical translation is therefore negative on this point: the study reports no dosimetric evaluation — no dose-volume histogram (DVH, the standard tool summarizing the dose received by the tumor and by each organ), no target coverage, no measured sparing of the parotids or spinal cord on the predicted CTs. A CT that "looks" more like the day's anatomy may still misinform the dose calculation, precisely because proton range is extremely sensitive to CT-number errors along the beam path. The 23.4% improvement in that error is the most clinically relevant result, but it remains an intermediate indicator, not proof of better dose.

What is good

A real clinical bottleneck, cleverly reframed. Online adaptive proton therapy is limited by the time and cost of day-of imaging. Flipping the problem — predicting anatomy before scanning, to prepare or triage upstream — is a fresh, concrete idea, not another variation on automatic segmentation. It targets a step where saving time has real clinical value.

A realistic "zero-shot" use. Reusing a pretrained registration model without patient-specific retraining is a pragmatic choice: no individual data to collect before it can be used, and it leverages experience accumulated across a population. This is the kind of constraint that decides, in practice, whether a method is deployable or stays a lab prototype.

A metric chosen with physical relevance. Reporting the CT-number error, not just an abstract image similarity, shows the authors know their field: in proton therapy, Hounsfield-unit accuracy directly conditions the proton-range calculation. Flagging that the benefit concentrates on high-change patients is also useful honesty: it delineates from the start who might benefit.

What is less good

A weak comparator. The chosen reference is the frozen planning CT — that is, "adapting nothing at all." That is the lowest possible bar. The real clinical alternatives — registering the patient's own prior images, or simply acquiring the day's scan — are not used as comparators. Beating no-adaptation does not prove you do better than current practice. This is the classic biased comparator failure mode, the one that inflates the impression of performance.

A surrogate metric, instead of the only one that counts. Image similarity and contour overlap are proxies. The currency of proton therapy is dose: target coverage, organ-at-risk sparing, robustness of the Bragg peak. No dosimetric evaluation is presented. This is the misleading-metric failure mode: a flattering number on an intermediate indicator can mask a null dosimetric benefit — or a negative one, if the predicted CT leads the dose calculation astray.

A single, small cohort and a fragile assumption. Eighty-eight patients, one center, a retrospective study, no external validation or prospective test, and the format of a workshop paper. Above all, the very principle — using another patient's trajectory as a prediction for this one — is exposed to population bias and individual variability: head-and-neck patients evolve idiosyncratically (which node shrinks, where weight is lost). Predicting the wrong change could be worse than predicting nothing, a risk that averaged metrics over 88 patients cannot rule out.

What it changes

For the research community, this is a stimulating proof of concept: transferring a "deformation trajectory" from one patient to another as an anatomical prior is a lead worth pursuing. The roadmap to make it credible is clear: add dosimetric endpoints (dose recomputation and DVH on the predicted CTs), an honest comparator (registration of the patient's own images), and external, multi-center validation, ideally prospective.

For clinicians, nothing to deploy today, and it must be said plainly: a predicted CT is a hypothesis about anatomy, not a substitute for the real day's image. Safety in proton therapy requires verifying the anatomy before delivering dose. One can imagine an upstream use down the line — triaging which patients will need re-planning, or preparing a "warm-start" plan to refine afterward — but never as a replacement for imaging verification.

For patients and the public, the reading lesson is useful: a model that "predicts your anatomy" from other patients is a planning aid, not a scan of you. And "22.8% better than doing nothing" does not mean "accurate." Gains measured on image similarity are not, in themselves, proof of safer or more effective treatment.

Further reading

The preprint is available on arXiv (2608.00831), submitted on 1 August 2026 and accepted at the Digital Twins for Healthcare (DT4H) workshop of the MICCAI 2026 conference. For context, see the physics of the Bragg peak in proton therapy, the principles of online and offline adaptive radiotherapy, deformable image registration and its evaluation, and foundation models for medical registration, along with the central question here: translating between geometric metrics and dosimetric endpoints.