Aneesh Rangnekar , Nikhil Mankuzhy , Jonas Willmann , Chloe Min Seo Choi , Abraham Wu , Maria Thor , Andreas Rimner , Harini Veeraraghavan
{"title":"Transformer-based cardiac substructure segmentation from contrast and non-contrast computed tomography for radiotherapy planning","authors":"Aneesh Rangnekar , Nikhil Mankuzhy , Jonas Willmann , Chloe Min Seo Choi , Abraham Wu , Maria Thor , Andreas Rimner , Harini Veeraraghavan","doi":"10.1016/j.phro.2026.101011","DOIUrl":null,"url":null,"abstract":"<div><h3>Background and Purpose</h3><div>Accurate segmentation of cardiac substructures on computed tomography (CT) scans is essential for radiotherapy planning. This study evaluated whether pretrained transformers enabled data-efficient training using a fixed architecture with balanced curriculum learning while achieving robust generalization to imaging and patient variations.</div></div><div><h3>Materials and Methods</h3><div>A hybrid pretrained transformer-convolutional network, self-distilled masked image transformer (SMIT), was fine-tuned using lung cancer patient scans (Cohort I, training N = 180) and tested on held-out Cohort I lung cancer scans (testing N = 60) and breast cancer scans (Cohort II, N = 65). Two configurations were evaluated: SMIT-Balanced (32 contrast-enhanced CTs, 32 non-contrast CTs) and SMIT-Oracle (180 CTs). Performance was compared with nnU-Net and TotalSegmentator. Segmentation accuracy was assessed primarily using the 95th percentile Hausdorff distance (HD95), along with radiation dose and overlap-based metrics as secondary endpoints.</div></div><div><h3>Results</h3><div>SMIT-Balanced approached SMIT-Oracle performance despite using 64% fewer training scans, with mean HD95 of 6.6 versus 5.4 mm in Cohort I and 10.0 versus 9.4 mm in Cohort II. On the Cohort I held-out test set, SMIT-Balanced mean HD95 was within 1.0 mm of nnU-Net. Cross-cohort testing showed larger accuracy degradation with nnU-Net than SMIT-Balanced (62% versus 50%, absolute change 4.5 mm versus 3.4 mm). Dose metrics derived from SMIT-Balanced were equivalent to manual delineations.</div></div><div><h3>Conclusions</h3><div>Balanced curriculum training reduced labeled data requirements within the SMIT architecture. SMIT-Balanced was comparable to nnU-Net on Cohort I held-out data and showed smaller cross-cohort HD95 degradation.</div></div>","PeriodicalId":36850,"journal":{"name":"Physics and Imaging in Radiation Oncology","volume":"39 ","pages":"Article 101011"},"PeriodicalIF":3.2000,"publicationDate":"2026-05-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Physics and Imaging in Radiation Oncology","FirstCategoryId":"1085","ListUrlMain":"https://www.sciencedirect.com/science/article/pii/S2405631626001119","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"2026/6/2 0:00:00","PubModel":"Epub","JCR":"Q2","JCRName":"ONCOLOGY","Score":null,"Total":0}
引用次数: 0
Abstract
Background and Purpose
Accurate segmentation of cardiac substructures on computed tomography (CT) scans is essential for radiotherapy planning. This study evaluated whether pretrained transformers enabled data-efficient training using a fixed architecture with balanced curriculum learning while achieving robust generalization to imaging and patient variations.
Materials and Methods
A hybrid pretrained transformer-convolutional network, self-distilled masked image transformer (SMIT), was fine-tuned using lung cancer patient scans (Cohort I, training N = 180) and tested on held-out Cohort I lung cancer scans (testing N = 60) and breast cancer scans (Cohort II, N = 65). Two configurations were evaluated: SMIT-Balanced (32 contrast-enhanced CTs, 32 non-contrast CTs) and SMIT-Oracle (180 CTs). Performance was compared with nnU-Net and TotalSegmentator. Segmentation accuracy was assessed primarily using the 95th percentile Hausdorff distance (HD95), along with radiation dose and overlap-based metrics as secondary endpoints.
Results
SMIT-Balanced approached SMIT-Oracle performance despite using 64% fewer training scans, with mean HD95 of 6.6 versus 5.4 mm in Cohort I and 10.0 versus 9.4 mm in Cohort II. On the Cohort I held-out test set, SMIT-Balanced mean HD95 was within 1.0 mm of nnU-Net. Cross-cohort testing showed larger accuracy degradation with nnU-Net than SMIT-Balanced (62% versus 50%, absolute change 4.5 mm versus 3.4 mm). Dose metrics derived from SMIT-Balanced were equivalent to manual delineations.
Conclusions
Balanced curriculum training reduced labeled data requirements within the SMIT architecture. SMIT-Balanced was comparable to nnU-Net on Cohort I held-out data and showed smaller cross-cohort HD95 degradation.