AI-assisted abdominal ultrasound: a randomised crossover trial measures 52 seconds saved and 39% less hand travel, for 48 values corrected out of 184
Ten authors from Queensland University of Technology and CQUniversity in Brisbane put 32 healthy volunteers through two upper abdominal ultrasound examinations each — one manual, one assisted by AI Abdomen 3.5, the on-cart software on Siemens' ACUSON Sequoia platform — in randomised order, each examination performed by a different sonographer, with a depth camera measuring hand movement at the console. AI assistance cut examination time by 52.4 seconds (−9%), keystrokes by 28% and hand travel distance by 39%, without shifting the composite NASA-TLX workload score significantly (−3.9 points; P = 0.17), although the mental demand and effort subscales fell. The number worth remembering is elsewhere: sonographers corrected 48 of the 184 measurements the AI generated, and only one of the 32 assisted examinations required no intervention at all.
The context
Ultrasound is the only imaging modality where the image does not exist until an operator makes it. A radiologist reading a CT receives a completed study; the sonographer holds the probe, optimises in real time, recognises the view, freezes, measures, labels, stores. That dependence on the operator is also the constraint of the job: work-related musculoskeletal disorders are highly prevalent among sonographers, and recent literature attributes them not only to mechanical factors — repetitive movement, static posture, sustained console interaction — but to cognitive and psychosocial load as well.
Manufacturers have responded with on-cart assistance software. These systems run on the live image stream, recognise the anatomical view on screen, apply a label to it and place measurement callipers to produce a preliminary biometric value that the operator accepts, adjusts or rejects. This is a distinct category of medical AI: most imaging AI interprets an already stored study, whereas this one acts during acquisition, under the same clinical uncertainty as the operator, and redistributes their work rather than replacing it.
Independent evidence is thin. In echocardiography, a 2025 study reported reduced scan time and console interaction. More importantly, the PROMETHEUS trial, published in NEJM AI in 2025, tested acquisition-assist AI on the fetal anomaly scan: median duration down from 19.7 to 11.4 minutes, composite NASA-TLX from 46.5 to 35.2. That is this paper's natural comparator — and, as we will see, it does not say the same thing.
The method
The design. Prospective randomised crossover trial run in January and February 2026 in a university ultrasound laboratory, approved by the QUT ethics committee on 15 December 2025 (no. 10419), reported per CONSORT 2025, its crossover-trial extension, and the DECIDE-AI checklist — the reporting guideline for early-stage clinical evaluation of AI-driven decision support. A crossover trial means each participant serves as their own control: the defence against the dominant source of variance in ultrasound, patient body habitus.
Operators and participants. Two accredited sonographers with 11 and 25 years of adult abdominal experience, neither of whom had used the system before, each trained by a single standardised vendor session plus one practice scan. Thirty-two healthy volunteers recruited by convenience from the university community: 20 women (62.5%), mean age 25.3 years (SD 8.9; range 18–48), mean BMI 22.6 (3.6; 16.4–31.3). Hold on to that last figure: this is a young, lean cohort.
Randomisation. Four sequences of 8 participants, with condition order (AI-first or manual-first) and sonographer order randomised and balanced independently. Each participant was scanned once by each sonographer, with the assisted and manual examinations performed by different operators, five minutes apart. Each sonographer therefore performed 16 assisted and 16 manual examinations. The allocation sequence was held by a research assistant and not disclosed before each scan. Blinding was impossible: the AI workflow is visually distinct on the console.
The system. AI Abdomen Release 3.5 (VB30), vendor-integrated software on an ACUSON Sequoia platform (Siemens Medical Solutions), 5C1 curvilinear transducer, standard abdominal preset. No training or tuning was performed by the team. Importantly: the algorithm class, training population and preclinical performance are not publicly disclosed by the manufacturer. The protocol covers longitudinal and transverse views of the pancreas, aorta, inferior vena cava, liver, gallbladder, common bile duct, kidneys and spleen, restricted to the measurements the AI supports. Doppler is excluded — it does not exist in the assisted workflow.
The measurement. This is the original contribution. An Intel RealSense D435 depth camera fixed above the console records at 848 × 480 and 30 frames per second. Sonographers wore dark blue nitrile gloves; thresholding in hue-saturation-value colour space isolates glove pixels, then a channel transformation shifts them towards the skin-tone profile expected by the MediaPipe hand detector, which returns 21 landmarks per frame. The trajectory is despiked and smoothed with a One Euro filter, yielding four metrics: cumulative hand travel distance, hover time over the keyboard zone, average jerk (the derivative of acceleration — an index of movement smoothness) and scan duration. Keystrokes were counted in real time by an independent observer. Perceived workload was measured after each examination with the weighted NASA-TLX, a six-dimension instrument (mental, physical and temporal demand, performance, effort, frustration) with a pairwise weighting procedure.
The analysis. Linear mixed-effects models — condition, sonographer and scan order as fixed effects, participant as a random intercept. Scan duration and weighted NASA-TLX were the two prespecified primary outcomes; keystrokes, travel, hover and jerk the secondary ones; everything else is exploratory, with no multiplicity adjustment. No a priori power calculation; post hoc, 32 pairs give 80% power to detect a 7.9-point NASA-TLX effect. The study was not registered in a clinical trials registry, the authors reasoning that the intervention did not target a health outcome.
The results
Efficiency. All five operational outcomes point the same way, all significant under both parametric and paired Wilcoxon testing. Scan duration: 592.4 s manual versus 540 s assisted, i.e. −52.4 s (95% CI −81.2 to −23.7; −9%; P = 0.001). Keystrokes: 197 → 142, i.e. −55 (−28%; P < 0.001). Hand travel: 11.81 → 7.24 m, i.e. −4.57 m (−39%; P < 0.001). Keyboard hover time: 64.4 → 31.1 s (−52%; P < 0.001). Average jerk: 3.06 → 2.23 m/s³ (−27%; P < 0.001). These gains reflect faster completion of the same protocol: 192/192 measurements obtained manually, 189/192 with assistance.
Workload. The second primary outcome is negative. Weighted NASA-TLX goes from 45.8 to 41.9, i.e. −3.9 points (95% CI −9.3 to +1.5; P = 0.17): the interval contains zero, and the study was underpowered for an effect this size. In subscale analysis — exploratory — mental demand falls by 6.3 points (P = 0.03) and effort by 7.0 points (P = 0.04), with the other four dimensions unchanged. The authors emphasise the point that interests them: no subscale increases, so no visible transfer of load into supervisory activity. Separately, scan duration is associated with perceived workload within participants (0.089 NASA-TLX points per second; 95% CI 0.060–0.118; P < 0.001).
Corrections. The AI produced a value for 184 of the 189 measurements obtained; in the remaining five (gallbladder wall ×2, common bile duct ×2, right kidney length ×1) it generated nothing and the operator measured manually. Sonographers modified 48 of the 184 values and corrected 33 anatomical labels. Only one of the 32 assisted examinations required no intervention. Modification rates vary sharply by measurement: liver span 18/31 (58%; 95% CI 41–74), gallbladder wall thickness 11/30 (37%; 22–55), common bile duct diameter 7/29 (24%; 12–42). The two operators corrected a similar overall proportion (21/90 and 27/94) but not the same measurements. Labelling errors cluster on adjacent, sonographically confusable structures in the epigastric region — the most frequent pair being pancreas corrected to inferior vena cava (n = 4) — then on contralateral organs, right kidney corrected to left kidney (n = 3).
Operational translation. Nine percent of duration saved on an abbreviated ten-minute examination is under a minute. Across a list of twenty abdominal ultrasounds, that would be roughly seventeen minutes — enough for one more examination, provided the gain survives in real patients and the saved time returns to the schedule rather than to the working day. Conversely, a 58% correction rate on liver span means that, for that particular measurement, the operator overrides the AI more than half the time: the proposed value is a starting point, not a result.
What is good
The movement counting is instrumented, not self-reported. Almost the entire literature on workload in imaging rests on questionnaires. Here four of the six metrics come from a depth camera with a fully described pipeline — glove segmentation in HSV space, channel transformation to fool a hand detector trained on bare hands, restriction to the central 40% of the field to eliminate screen reflections, 0.3 m per frame despiking, One Euro filter — and the tracking software is published on GitHub. Keystrokes are counted separately by a human observer, providing a cross-check of the same phenomenon through an independent channel. This is a level of methodological detail rarely seen for this kind of measurement.
The crossover design is used for what it is worth, and the robustness analyses follow. Each participant is their own control, condition order and operator order are randomised and balanced separately, and the authors go looking for residual confounding: models with a condition × sonographer interaction term, a study-sequence covariate, leave-one-participant-out re-fits across all 32 participants, paired non-parametric tests as a check. They find a trend that is awkward for them — perceived workload declines across the study period in both conditions (−0.93 points per participant assisted, P < 0.001; −0.55 manual, P = 0.04) — identify it as general familiarisation, and note that the crossover protects the contrast but not the absolute magnitude of the cognitive effect.
The negative result is not repainted, and the correction inventory is published. Composite NASA-TLX was a prespecified primary outcome; it is non-significant, and the abstract says so before mentioning the subscales, noting that those are exploratory and unadjusted for multiplicity. The authors' conclusion — that the gains are "better read as a reshaping of operator work than its removal" — is more cautious than their own numbers would require. As for the corrections, they are given measurement by measurement, with Wilson intervals, a breakdown by operator, and a complete table of corrected label pairs: precisely the raw material a department would need to decide which functions to enable and which to switch off.
What is less good
Population bias here is close to maximal. Mean BMI 22.6, mean age 25.3, healthy volunteers with no abdominal pathology, fasted, abbreviated protocol without Doppler. Yet the difficulty of abdominal ultrasound comes precisely from what was excluded: excess weight, bowel gas, anatomy remodelled by surgery, pathology to characterise. A view-recognition model trained on undisclosed data could well see its reliability collapse in an obese patient — and that is exactly where the time saving would matter most. The authors acknowledge this explicitly, but the gap between the tested population and the target population remains the main limit on extrapolation.
The rate of uncaught errors is structurally unknown. The protocol has no independent reference standard: sonographer verification is the reference process. Every recorded modification is therefore an AI error the operator caught, and the design can say nothing about the ones they did not. This is the automation bias failure mode — the documented tendency of an operator to accept a plausible machine suggestion — and two sonographers with 11 and 25 years of experience are the most resistant profile there is; the paper itself notes that a less experienced operator would be more exposed. Add the complete absence of any assessment of image quality and measurement accuracy, which the authors defer to separate work: a quality-for-speed trade-off cannot be excluded. The PROMETHEUS comparison is instructive here — in that trial, the quality of AI-selected images was initially rated below manually acquired images, and that shortfall was visible only because an independent quality assessment had been planned, which is not the case here.
Two operators, one machine, and fragile significance on the subscales. With only two sonographers, condition and operator are partially confounded at participant level; the authors handle this with balancing and interaction models, but acknowledge that the time saving and the NASA-TLX effect are concentrated in one of the two. An effect that half the operator population does not enjoy is not generalisable as it stands. Furthermore, six subscales were tested at α = 0.05 with no multiplicity adjustment, and the two positive results come in at P = 0.03 and P = 0.04: under elementary Bonferroni correction, neither would survive. Finally, one system and one manufacturer: nothing says the results transport to another platform. On interests, the declaration is candid — Siemens Healthineers provided access to the software evaluated, with no role in design, analysis or the decision to submit, and the funding (a QUT vacation research scheme grant, and the Australasian Sonographers Association for participant travel costs) is modest and free of apparent conflict. Still, the algorithm evaluated is a commercial black box whose class, training data and preclinical performance are all non-public.
What it changes
For the research community. The paper delivers two reusable things. The first is methodological: an instrumented, open apparatus for objectively measuring operator-console interaction, in a field that had been making do with self-reported NASA-TLX. The second is an empirical counterpoint to PROMETHEUS. Both studies evaluate acquisition-assist AI; one finds a clear drop in composite workload, the other does not. The authors offer an operational explanation rather than a disagreement: PROMETHEUS automated the freeze-save-measure cycle across 13 standard planes and captured silently during acquisition, with human review deferred to the end of the examination; the system tested here displays labels and callipers on the live image and asks for validation concurrent with acquisition. If that reading is right, then when you ask the human to verify matters as much as how well the model performs — a testable hypothesis, and one that extends well beyond ultrasound.
For clinicians. Nothing changes today: a non-peer-reviewed preprint, on healthy volunteers, with no image quality assessment, and the medRxiv header explicitly states these results should not guide practice. The useful reading lies elsewhere. First, the measured efficiency gains are real but modest, and concentrated in one operator; nobody should buy this class of software on a throughput promise without a local measurement. Second, and more importantly, the distribution of corrections tells you where the tool is reliable and where it is not: accepting an AI-proposed liver span outright would be imprudent when it is overridden 58% of the time, whereas other measurements almost never are. A sensible deployment enables the reliable functions and disables the rest, rather than enabling the bundle.
For patients and the public. This study is not about diagnosis but about working conditions, which is what makes it interesting. An AI that cuts by 39% the distance a hand travels across a console acts on a documented physical burden in a profession where musculoskeletal disorders are endemic — but the authors are clear that their metrics cover only the keyboard hand, not the posture of the arm holding the probe nor the force applied, and that nothing here should be read as evidence of reduced injury risk. As for the time saved, it will reach patients — as shorter waits or extra examinations — only if three conditions hold: that the gain persists in clinically representative examinations, that image quality and measurement accuracy do not degrade, and that the saved time is returned to the schedule. None of the three is established.
Further reading
- The paper: Quantifying Human–AI Workflow in Abdominal Ultrasound: A Prospective Randomised Crossover Study, Hsiao, Clifford, Lin et al., medRxiv 2026.08.17.26360254, CC BY 4.0. Preprint, not peer reviewed. Funding: QUT Vacation Research Experience Scheme and the Australasian Sonographers Association; Siemens Healthineers provided access to the evaluated software.
- The hand-tracking system, released open source: HTS — Hand Tracking System, Lin, 2026. Data and R analysis code under mediated access: QUT Research Data Finder.
- The closest comparator: Day et al., AI to Assist in the Fetal Anomaly Ultrasound Scan: A Randomized Controlled Trial, NEJM AI 2:AIoa2400747, 2025.
- On automation bias in imaging: Dratsch et al., Automation Bias in Mammography, Radiology 307:e222176, 2023. On the reporting guideline used: Vasey et al., DECIDE-AI, BMJ 377:e070904, 2022.
- On Tatakoto: agentic AI on thyroid ultrasound across 35 centres, automatic segmentation of the operative workflow in colorectal surgery, and what the choice of comparator makes a prospective evaluation say.