AI contouring for radiotherapy cut delineation time by 59% in a single-centre study
A deep-learning system drew tumour targets and organs at risk automatically. It saved time and improved consistency — but oncologists still reviewed and edited the contours before treatment.
| Group | Value (value) |
|---|---|
| Pre-implementation | 0.8 |
| Post-implementation | 0.88 |
| After update | 0.95 |
Before a course of radiotherapy can be planned, someone has to draw, slice by slice on a CT scan, the exact outline of the tumour target and of the healthy "organs at risk" the beam must spare. This delineation — contouring — is slow, is done by hand, and varies from one clinician to the next. A study in the Journal of Applied Clinical Medical Physics, published online on 28 August, followed one hospital as it handed that job to a deep-learning system, and found it cut contouring time by roughly 59% while improving how consistently the contours were drawn — with the important caveat that oncologists still reviewed and revised the output before it reached a patient [s1].
What was measured
This was a longitudinal analysis of 150 patients treated for rectal cancer, split into three cohorts of 50: before the automatic contouring system was introduced, after it was introduced, and after it received a software update [s1]. The team compared the system's unedited automatic contours against the final contours that were actually used for treatment, using the Dice similarity coefficient (DSC), a 0-to-1 overlap score where 1 is a perfect match [s1]. Six oncologists also contoured 21 additional cases three ways — fully manually, with the first-generation system (Auto1), and with the updated system (Auto2) — to measure time, consistency and accuracy, and two senior oncologists blindly rated the clinical acceptability of the contours on a five-point scale [s1].
The update mattered more than the launch
For the clinical target volume — the tumour region — the automatic contours already matched the final ones closely, with a mean DSC of 0.87 ± 0.04 before implementation and 0.88 ± 0.04 after, a difference that was not statistically significant (P = 0.067) [s1]. The gains were larger for organs at risk, where mean DSC rose from 0.80 ± 0.06 to 0.88 ± 0.05 (P < 0.001) [s1].
The bigger jump came with the update. After it, mean DSC improved to 0.93 ± 0.04 for the target volume and to 0.95 ± 0.02 for organs at risk (both P < 0.001), and the mean rate at which contours failed and had to be redone fell by about 80.6% [s1]. On the 21-case comparison, the updated system cut total contouring time by about 58.8% versus fully manual work and by about 21.9% versus the first-generation system, while delivering the best inter-observer consistency (0.95 ± 0.03) and accuracy (0.94 ± 0.03) for the target volume [s1].
That an update produced a step-change is the study's most useful message. An AI contouring tool is not a fixed instrument; its performance depends on the version running, and a hospital that validated the first release cannot assume the numbers still hold after an upgrade — or that they would be the same on a different scanner, a different cancer, or a different patient mix.
The human stayed in the loop
None of these contours went to a linear accelerator untouched. In the blinded review, 99.2% (125 of 126) of the oncologist-revised final contours scored 4 or higher on the five-point acceptability scale [s1]. The system's unedited output was rated lower, though the updated version was clearly better than the first: 4.02 ± 0.25 for Auto2 against 3.26 ± 0.49 for Auto1 (P < 0.001) [s1]. In other words, the automatic contours were a strong starting point that a clinician corrected, not a finished plan — which is exactly how the study's authors frame the tool, as guidance that reduced workload and variation rather than a replacement for review [s1].
Limits worth stating plainly
This is a single-centre, largely retrospective study of one cancer site and one commercial system, with the geometric comparison resting on 150 patients and the time-and-acceptability comparison on 21 cases contoured by six clinicians [s1]. A DSC is a measure of overlap, not of clinical safety: two contours can score well on Dice and still differ in the one place that matters dosimetrically. The study did not report patient outcomes, and it cannot tell you whether the time saved changed anything a patient would notice. It joins a wider pattern in medical AI, where devices are cleared and adopted well ahead of evidence that they change patient outcomes, and where regulators' own testing of AI imaging tools has been found wanting.
There is also an automation-complacency risk that this design cannot see. When a tool is usually right, reviewers can stop scrutinising it — the same trap documented for over-frequent clinical alerts that clinicians learn to dismiss. A recent systematic review found the field still lacks even an agreed way to measure that kind of drift in how carefully clinicians respond to decision-support output [s2]. A contour that is edited less because it is trusted more is a benefit only if the trust is earned on every case.
What to watch
Whether auto-contouring evaluations move beyond overlap scores to dosimetric and patient outcomes, and whether vendors and hospitals report performance per software version — so that an upgrade is revalidated rather than assumed. This matters most for cancers managed with a mix of surgery, surveillance and radiotherapy, and for systems where radiotherapy capacity is already stretched, where any real time saving would go furthest.
Sources
- [s1] Real-world clinical impact of implementing and updating a deep learning-based automatic contouring system in rectal cancer radiotherapy. Journal of Applied Clinical Medical Physics, online 28 August 2026. https://doi.org/10.1002/acm2.70768
- [s2] Alert fatigue measurement in clinical decision support: a systematic review. Journal of the American Medical Informatics Association, 18 May 2026. https://doi.org/10.1093/jamia/ocag064
Sources
- Real-world clinical impact of implementing and updating a deep learning-based automatic contouring system in rectal cancer radiotherapy — Journal of Applied Clinical Medical Physics , August 28, 2026
- Alert fatigue measurement in clinical decision support: a systematic review — Journal of the American Medical Informatics Association , May 18, 2026
AI can flag pancreatic cancer on ordinary CT scans, but only in retrospective tests
A deep-learning model reads the pancreas on non-contrast CT taken for other reasons. It reached very high accuracy in large tests — none of them a prospective screening trial.
AI for pulmonary embolism on CT buys speed, not accuracy — three 2026 studies
Where AI clot-detection was tested head-to-head, it still missed more emboli than radiology trainees, especially small peripheral ones. Its clearest win was cutting time to diagnosis.
An AI out-read experts on ovarian ultrasound across 20 centres in eight countries
A transformer model trained on 17,119 scans beat expert and non-expert examiners on every metric and cut simulated referrals by 63%. The catch is who built it and the validation the rest of the field still lacks.
The FDA's blueprint for AI medical devices: what a maker must show, cradle to grave
A January 2025 draft guidance lays out the documentation the agency expects across an AI device's whole life, including bias checks across demographic groups and postmarket monitoring. It is not yet binding.