SynthRAD2025 Challenge: sCT Synthesis
- SynthRAD2025 is a benchmark that enables CT synthesis from MRI and CBCT data for MRI-only radiotherapy and adaptive workflows.
- The challenge assesses image quality using metrics like HU fidelity, geometric consistency, segmentation Dice scores, and dosimetric accuracy.
- The dataset, collected from five European hospitals, features diverse imaging protocols and a standardized preprocessing pipeline for clinical validation.
SynthRAD2025 is a Grand Challenge for generating synthetic computed tomography (sCT) for radiotherapy from magnetic resonance imaging (MRI) and cone-beam computed tomography (CBCT). It provides a large multi-center benchmark for MRI-to-CT and CBCT-to-CT synthesis, with evaluation spanning image similarity, geometric consistency, and dosimetric accuracy for photon and proton plans. Building on SynthRAD2023, the benchmark covers 2,362 patients from five European university hospitals and is intended to support MRI-only radiotherapy, MR-guided photon and proton radiotherapy, CBCT-based dose calculations, and adaptive radiotherapy workflows (Thummerer et al., 24 Feb 2025, Rogowski et al., 13 May 2026).
1. Clinical scope and benchmark tasks
The challenge is organized into two benchmark tasks. Task 1 targets MRI-to-CT synthesis, motivated by the absence of electron-density information in MR images. Task 2 targets CBCT-to-CT translation, motivated by CBCT artifacts and scatter that limit accurate dose calculations. In both tasks, the goal is to produce CT-equivalent images with reliable Hounsfield Unit (HU) values (Thummerer et al., 24 Feb 2025).
| Task | Paired data | Anatomical regions |
|---|---|---|
| MRI-to-CT | 890 paired T1-weighted MR and planning CT images | Head-and-neck, thorax, abdomen |
| CBCT-to-CT | 1,472 paired cone-beam CT and planning CT images | Head-and-neck, thorax, abdomen |
The full cohort was assembled from UMC Groningen, UMC Utrecht, Radboud UMC, LMU University Hospital Munich, and University Hospital of Cologne. Each center contributed cases for all CBCT-to-CT subsets. MR-to-CT subsets were available at four centers for head-and-neck, at two centers for thorax, and at three centers for abdomen. This design makes heterogeneity a central property of the benchmark rather than a nuisance variable (Thummerer et al., 24 Feb 2025).
The challenge report frames the benchmark in explicitly clinical terms. CT remains fundamental for treatment planning because it provides electron-density information for dose calculation; repeated CT acquisition adds radiation exposure and logistical burden; MRI offers superior soft-tissue contrast but lacks electron-density information; and CBCT enables daily imaging at the treatment unit but suffers from scatter and artifacts that degrade HU accuracy. Within that context, sCT generation is positioned as a route to MRI-only planning and CBCT-based adaptive workflows without extra CT scans (Rogowski et al., 13 May 2026).
2. Dataset composition and acquisition heterogeneity
A defining feature of SynthRAD2025 is acquisition diversity. MR data span a ViewRay 0.35 T MR-Linac with balanced steady-state free-precession sequences, 1.5 T and 3 T Philips Ingenia systems, and Siemens Skyra and Prisma systems using T1-weighted turbo-gradient-echo Dixon, radio-frequency-spoiled gradient-echo, or turbo spin-echo sequences. Reported voxel spacings range from mm to mm (Thummerer et al., 24 Feb 2025).
Planning CTs were acquired on Philips Big Bore and Biograph, Siemens Somatom Definition AS, Toshiba Aquilion and LB, and GE Optima platforms, with kV settings of 90–140 kV, tube currents from 19 to 534 mA, slice thicknesses of 1–5 mm, and reconstruction diameters spanning 320–800 mm. CBCT data originated from Elekta XVI v5.x, Varian TrueBeam OBI, and IBA Proteus P+ systems under clinically relevant protocols with 100–140 kVp, 10–320 mA, slice thickness 1–2.5 mm, and pixel spacings 0.5–2.0 mm (Thummerer et al., 24 Feb 2025).
This multi-vendor, multi-protocol structure captures variability in hardware, field strengths, immobilization devices, and patient positioning. A plausible implication is that benchmark performance is constrained not only by synthesis fidelity in a narrow paired setting, but also by robustness across institution-specific acquisition regimes. The challenge report is consistent with that interpretation: it emphasizes multi-center, multi-anatomy generalization and later identifies center-wise differences and anatomy-specific variability in the final results (Rogowski et al., 13 May 2026).
All imaging data is provided in MetaImage format, and metadata including manufacturer, model, kVp, sequence name, TE/TR, voxel spacing, and registration parameters are collated into structured CSV or XLSX files. The dataset is accessible under the SynthRAD2025 collection at https://doi.org/10.5281/zenodo.14918089 (Thummerer et al., 24 Feb 2025).
3. Preprocessing, masking, and quality assurance
SynthRAD2025 distributes data after a fully automated, publicly available preprocessing pipeline hosted at https://github.com/SynthRAD2025/preprocessing. The pipeline performs rigid registration of MRIs or CBCTs to planning CTs using Elastix parameter files optimized per anatomy; automated defacing of CT and MRI via TotalSegmentator v2.3.0; resampling to an isotropic in-plane spacing of mm and slice thickness of 3 mm; automatic patient-outline mask generation through histogram-based thresholding followed by morphological erosion and dilation; dilation of the mask to include surrounding air; cropping to a ten-pixel margin around the dilated outline; and conversion and compression into MetaImage format with INT16 pixel representation. For validation and test only, deformable CTMR/CBCT registration via Elastix generates geometry-matched ground-truth CTs (Thummerer et al., 24 Feb 2025).
The defacing procedure is anatomically explicit. The skull and brain are auto-segmented, anterior landmarks are defined on the central sagittal slice, a facial bounding box is generated, and voxels anterior to this box are overwritten with background intensities: HU for CT and $0$ for MR. This is paired with body-contour masking so that evaluation is limited to relevant anatomy rather than empty background (Thummerer et al., 24 Feb 2025, Rogowski et al., 13 May 2026).
Quality assurance is human-centered rather than purely algorithmic. Each center performed visual quality checks on defaced images to confirm complete facial anonymization. Across the dataset, three-plane overview images overlay input and CT volumes together with the patient outline mask, enabling inspection of rigid-registration alignment, major imaging artifacts, and mask completeness. Selection for validation and test additionally required visually acceptable deformable registration quality and, where available, satisfactory organ-at-risk or target structure alignment. No automated numeric image-quality metrics such as mutual information scores were applied; clinical usability was guaranteed by expert review (Thummerer et al., 24 Feb 2025).
The split strategy is fixed at 65% training, 10% validation, and 25% test, applied per task and anatomy with minor center-specific adjustments. Training data were released on 01 March 2025 with input images, ct.mha, and mask.mha. Validation input and masks were released on 01 June 2025, while ct.mha and ct_def.mha were withheld until challenge end. The test set is scheduled for full modality release on 01 March 2030. Deformed CTs are withheld in the training split to prevent paired-image-registration bias (Thummerer et al., 24 Feb 2025).
4. Evaluation protocol and metric structure
SynthRAD2025 evaluates submissions on validation and test sets using image-based, segmentation-based, and dose-based criteria. Participants submit sCT volumes in .mha format. The dataset paper specifies representative formulas for Mean Absolute Error, Root Mean Squared Error, Structural Similarity Index, dose-difference metrics, and gamma analysis, while the challenge report formalizes the test-cohort reporting in terms of MAE, PSNR, MS-SSIM, multi-class Dice, HD95, and pass rates for photon and proton plans (Thummerer et al., 24 Feb 2025, Rogowski et al., 13 May 2026).
Within the body contour , the Mean Absolute Error is reported as
0
The challenge report also defines multi-class Dice for 1 segmented structures as
2
For dosimetry, the dataset paper gives the gamma formulation
3
and the challenge report specifies 4 and 5 mm, with gamma pass rates evaluated in the volume receiving at least 6 of the prescribed dose (Thummerer et al., 24 Feb 2025, Rogowski et al., 13 May 2026).
This metric suite is intentionally heterogeneous. Image metrics quantify HU fidelity and structural similarity; segmentation metrics quantify geometric agreement of derived structures; and dose metrics test whether residual synthesis errors remain acceptable under downstream treatment-planning calculations. The challenge report later shows that these metric families are correlated, but not interchangeable, which is central to interpreting leaderboard performance (Rogowski et al., 13 May 2026).
5. Methodological landscape and representative challenge entries
No official baseline algorithms are provided. Participants commonly adapt deep neural networks such as U-Net, GAN-based models, or registration-augmented synthesis pipelines (Thummerer et al., 24 Feb 2025). In the final challenge report, the top-ranked teams predominantly used CNN encoder-decoder networks such as U-Net, nnU-Net, and ResUNet; PatchGAN discriminators and multi-head segmentation-aware discriminators; and 2.5D or 3D patch-based variants to balance context and GPU memory. Flow-matching and denoising diffusion models were represented, but on average performed below CNN/GAN approaches in this challenge. All teams used supervised learning with paired deformably registered CT references, and common loss components included MAE, MSE, perceptual losses, adversarial losses, feature-matching losses, and segmentation-aware or anatomical feature-prioritized losses using pretrained segmenters such as TotalSegmentator (Rogowski et al., 13 May 2026).
A representative generative submission is the fully 3D flow-matching framework described in "Flow Matching for Conditional MRI-CT and CBCT-CT Image Synthesis" (Hadzic et al., 6 Oct 2025). That model follows the probability-flow ODE formulation of Lipman et al. (2022), conditions on features extracted from MRI or CBCT by a lightweight 3D encoder with two 7 convolutional layers, and predicts a 3D velocity field using a 3D U-Net with four resolution levels, channel sizes 8, self-attention at the 9 and 0 resolutions, dropout 1, and approximately 41.2 million parameters. Volumes are resampled to 2 voxels at 3 isotropic resolution; optimization uses AdamW with learning rate 4, weight decay 5, and effective batch size 6; and inference solves 7 with a 4th-order Runge–Kutta integrator using 32 steps. Training used an 80 GB NVIDIA A100, inference used a GeForce RTX3090, and inference time was approximately 2 minutes per volume (Hadzic et al., 6 Oct 2025).
That submission trained separate models for Task 1 and Task 2 across abdomen, head and neck, and thorax. For Task 1, the training set comprised 578 paired MR–CT volumes, distributed as 175 abdomen, 221 head and neck, and 182 thorax, with 89 unpaired MR scans for validation and 223 MR scans in the test set. For Task 2, the training subset was reported as 309 abdomen, 325 head and neck, and 321 thorax paired CBCT–CT volumes, with 148 unpaired CBCT scans for validation and 369 CBCT scans in the test set. Validation results were 8 HU MAE, 9 dB PSNR, 0 MS-SSIM, 1 mDice, and 2 mm HD3 for MRI4sCT, and 5 HU MAE, 6 dB PSNR, 7 MS-SSIM, 8 mDice, and 9 mm HD0 for CBCT1sCT. Test-phase metrics were not yet public in that paper. The authors reported recovery of global anatomy without obvious artifacts, but also blurriness and loss of fine details, attributed primarily to the uniform 2 training resolution imposed by memory and runtime constraints (Hadzic et al., 6 Oct 2025).
6. Reported outcomes, correlations, and unresolved issues
The challenge report summarizes 803 participants and 12 out of 13 valid submissions. For Task 1, the top submission was KoalAI with 3 HU MAE, 4 dB PSNR, 5 MS-SSIM, 6 mDice, 7 mm HD8, photon 9, and proton 0. For Task 2, the top submission was MixCT with 1 HU MAE, 2 dB PSNR, 3 MS-SSIM, 4 mDice, 5 mm HD6, photon 7, and proton 8 (Rogowski et al., 13 May 2026).
Several result patterns are explicit. Top CBCT-to-CT submissions outperformed MRI-to-CT on all reported metrics, suggesting a smaller domain gap and more robust HU mapping. Photon plans consistently achieved at least 9 gamma pass rates among top methods, while proton plans remained more sensitive, with approximately 0 in Task 1 and approximately 1 in Task 2 for the leading methods. Head-and-neck cases yielded the highest and most consistent metrics across centers, whereas thorax and abdomen exhibited greater inter-patient and inter-center variability due to respiratory motion, scatter artifacts, and non-uniform CBCT field of view. Center-wise differences were statistically significant in some regions, and proton metrics showed larger center effects than photon metrics (Rogowski et al., 13 May 2026).
The report also quantifies inter-metric relationships through Spearman correlations. MAE versus PSNR reached 2 in Task 1 and 3 in Task 2; MAE versus MS-SSIM reached 4 and 5; and MS-SSIM versus Dice reached 6 and 7. By contrast, photon gamma versus MAE was 8 in Task 1 and 9 in Task 2, proton gamma versus MAE was $0$0 and $0$1, and DVH-based metrics showed near-zero correlation $0$2 with other metrics. This directly supports the conclusion that image quality is insufficient as a dosimetric surrogate (Rogowski et al., 13 May 2026).
Residual failure modes are localized rather than uniformly distributed. The report states that residual errors concentrated at air–tissue and bone–soft-tissue interfaces propagate into proton dose discrepancies more strongly than photon dose discrepancies. It also notes that paired training depends on deformable registration, so residual misregistration introduces irreducible error. These observations structure the challenge’s forward-looking recommendations: incorporate dose-aware objectives during model training, improve registration quality or consider slice-cohort acquisitions, extend robustness to motion and out-of-distribution anatomies, and standardize multi-metric evaluation including DVH parameters and gamma pass rates before clinical deployment (Rogowski et al., 13 May 2026).
Taken together, SynthRAD2025 establishes a benchmark in which dataset heterogeneity, clinically grounded preprocessing, and dose-based validation are inseparable from model ranking. The principal lesson of the published results is not merely that deep learning can produce CT-like images, but that clinically relevant validation for sCT generation requires simultaneous attention to HU fidelity, structural agreement, and downstream dose behavior, especially for proton therapy (Rogowski et al., 13 May 2026).