ESTRO 2026 Congress Report | Physics track I Proffered Papers Session

Among the work shared at ESTRO 2026, one Swiss-Danish collaboration tackled a problem that quietly constrains almost every AI- and trial-led design effort in the field: the chronic shortage of large, diverse and complete multimodal patient datasets.

Modern head-and-neck (H&N) radiotherapy relies on several concurrent imaging modalities. Computed tomography (CT) provides the electron density needed for dose calculation; T1- and T2-weighted magnetic resonance imaging (T1w- and T2w-MRI) give the soft-tissue contrast used for tumour delineation; and 18F-fluorodeoxyglucose positron emission tomography (FDG-PET) reflects metabolic activity for biological target definition. Yet datasets containing all four modalities, co-registered for the same patient, are scarce, single-institutional and privacy-restricted. At the same time, the randomised controlled trial (RCT), as radiation oncology’s gold standard, is slow and costly, often reporting only after the technology under evaluation has already moved on.

Virtual clinical trials (VCTs), in which simulations are populated by realistic “virtual patients”, have been proposed as a way to complement real cohorts and to run experiments that real trials cannot. Their Achilles’ heel has always been the virtual patient itself: traditional digital phantoms derive from a handful of manually segmented templates, while recent generative-AI models can synthesise only one modality at a time.

Addressing this gap was the focus of a collaborative study between the Centre for Proton Therapy at the Paul Scherrer Institute and ETH Zurich, Switzerland, and the Danish Centre for Particle Therapy together with Aarhus University and Aarhus University Hospital, Denmark. The work, presented by Mr. Muheng Li (Paul Scherrer Institute and ETH Zurich, Switzerland) on behalf of the cross-institutional team, is, to the authors’ knowledge, the first generative engine to jointly synthesise co-registered CT, T1w-MRI, T2w-MRI and FDG-PET 3 phantoms of the head and neck within a single model. The Swiss group contributed the generative-modelling expertise, while the Danish partners provided the curated four-modality clinical cohort and the downstream segmentation pipelines, a pairing of strengths that made the project possible.

How it works

A pre-trained medical variational autoencoder (VAE) first compresses each four-modality study into a compact 16-channel latent code. A guided-diffusion 3D U-Net is then trained in this latent space as a latent rectified flow, a technique that “straightens” the path between random noise and data, allowing high-quality samples to be drawn in as few as eight deterministic steps. Learning a four-modality virtual patient distribution from 609 H&N cancer patients treated at Aarhus University Hospital, Denmark, the engine generates a complete four-modality phantom at 1 mm isotropic resolution in roughly 12 seconds on a single GPU, about an order of magnitude faster than conventional diffusion samplers at matched fidelity.

Does it hold up?

The synthetic cohort held up well under scrutiny. Per-modality Fréchet Inception Distances (FID; lower is better) ranged from 16.3 for PET to 33.5 for CT, comparable to published single-modality 3D generators. To test whether the four modalities of each synthetic patient carry consistent tumour information, three independently trained nnU-Net segmenters delineated gross tumour volumes from different modality combinations. Their median agreement on synthetic phantoms (Dice 0.82–0.87) came within 0.02–0.04 of real data (0.86–0.88). The authors were refreshingly candid that a tail of roughly 20% of cases dragged the mean lower (0.65–0.69 versus 0.76–0.79 on real data), an effect they trace to the inherited, rather than purpose-built, latent space and to the small size of the tumour target.

Why it matters

The take-home message is that unified four-modality 3D synthesis at clinically relevant resolution is feasible today. Such an engine can, in principle, turn 609 real patients into an essentially unbounded supply of distinct virtual patients, reusable infrastructure for VCTs, AI generalisation studies and treatment-planning-system validation. This echoes priorities flagged in the recent European Organisation for Research and Treatment of Cancer (EORTC) State-of-Science report and in the VCT roadmap by Corinne Faivre-Finn and colleagues.

The full team: Mr. Muheng Li (Paul Scherrer Institute and ETH Zurich, Switzerland), Dr Jintao Ren (Aarhus University and Aarhus University Hospital, Denmark), Prof Antony J. Lomax (Paul Scherrer Institute and ETH Zurich, Switzerland), Prof Stine Sofia Korreman (Aarhus University and Aarhus University Hospital, Denmark), and corresponding author Dr Ye Zhang (Paul Scherrer Institute, Switzerland), frames the study as a proof of concept. Conditional generation on demographic and clinical covariates, the addition of dose distributions and four-dimensional motion, and multi-institutional training and validation are named as the next steps on the road towards regulatory acceptance.

Can you spot the fake patient?

Even if, for the most part, three expert segmentation networks cannot tell synthetic and real patients apart, can you? The authors invite the ESTRO community to put the phantoms to the test in an interactive online challenge: you view anonymised H&N images across the four modalities and guess whether each patient is real or AI-generated. The collected responses will help gauge how convincing the synthetic cohort really is to human experts. Try it here:

[CLICK HERE]

Author details

Muheng Li: Centre for Proton Therapy, Paul Scherrer Institute, Villigen PSI, Switzerland; Department of Physics, ETH Zurich, Zurich, Switzerland.

Email: muheng.li@psi.ch

Social media: https://www.linkedin.com/in/muheng-li-ba4094287/