---
title: 'CaloDiT-2: Transformer Diffusion for Fast Calorimeter Simulation'
url: https://www.emergentmind.com/topics/calodit-2
type: topic
---

# CaloDiT-2: Transformer Diffusion for Fast Calorimeter Simulation

CaloDiT-2 is a transformer-based diffusion model for fast simulation of electromagnetic calorimeter showers in collider experiments, introduced in "A Generalisable Generative Model for Multi-Detector Calorimeter Simulation" [2509.07700]. It is designed as a generative surrogate for Geant4 calorimeter simulation, with an emphasis not only on speed and fidelity but also on cross-detector generalisation. The model is formulated around detector-agnostic data representations, transformer diffusion blocks, and a pre-training-plus-adaptation regime on multiple detector geometries. In the reported experiments, it produces showers in approximately \(100\) ms on CPU and approximately \(3\) ms on GPU, matches Geant4 on critical shower observables, and adapts to novel detectors with substantially reduced data and training requirements [2509.07700].

## 1. Research setting and stated objective

High-energy physics experiments, including those at the Large Hadron Collider, rely on the Geant4 toolkit for high-fidelity calorimeter shower simulation, but the CPU cost is described as seconds per shower and becomes prohibitive at high luminosity [2509.07700]. CaloDiT-2 is positioned as a response to this computational bottleneck: a diffusion model using transformer blocks that directly emulates detector responses while preserving the observables that matter for calorimeter simulation.

The reported objective is broader than accelerating a single fixed detector geometry. CaloDiT-2 can be applied to specific geometries, as is the case for other models explored for the task, but its stated strength lies in generalisation across detectors through pre-training on multiple geometries and rapid adaptation to new ones. The work frames this as an ambition toward a "foundation model" for FastSim: train once on varied geometries, then adapt cheaply to novel detectors under design [2509.07700].

The scope of the presented model is specifically electromagnetic calorimeter shower generation. The conditioning variables are defined for incident photons with energy \(E \in [1,1000]\) GeV and angular variables \(\phi \in [0,2\pi]\) and \(\theta \in [0.87,2.27]\). A plausible implication is that the reported generalisation claims are anchored to this simulation regime rather than to calorimeter simulation in full generality.

## 2. Detector-agnostic representation and conditioning interface

A central component of CaloDiT-2 is a detector-agnostic data representation built from a virtual cylindrical scoring mesh aligned with the incident particle, exploiting approximate azimuthal symmetry [2509.07700]. In the Par04/LEMURS setup, voxels are defined in \((r,\phi,z)\) with fixed resolution \(9\times16\times45\). This mesh is independent of any real read-out geometry: all energy deposits from Geant4 are scored into the mesh, yielding a tensor \(x \in \mathbb{R}^{9\times16\times45}\).

This representation separates the learning problem from any one detector segmentation scheme. Instead of binding the network to a specific read-out layout, the model operates on a common volumetric coordinate system into which multiple detectors can be projected. That design is the basis for the multi-detector pre-training strategy.

The conditioning variables are the incident photon energy \(E\), the particle angles \(\phi\) and \(\theta\), and a detector identifier \(G \in \mathrm{one\text{-}hot}(K+1)\), where \(K\) is the number of pre-training detectors and the extra bit flags an "unknown" geometry. Voxel preprocessing is given by
$$
\hat x_i = \frac{\log(x_i + \epsilon) - \mu}{2\sigma},
$$
where \(\mu\) and \(\sigma\) are the mean and standard deviation of log-energies over the entire pre-training set, and \(\epsilon = 10^{-6}\). The scalar conditions are normalised as \(E/E_{\max}\) and \(\theta/\pi\), while \(\phi\) is encoded as \((\sin\phi,\cos\phi)\) to preserve cyclicity [2509.07700].

The use of an "unknown" geometry flag is notable because it makes the conditioning scheme explicitly compatible with adaptation scenarios. This suggests that the representation is intended not merely for interpolation among known detectors but also for transfer to unseen geometries.

## 3. Diffusion formulation and consistency distillation

CaloDiT-2 follows Karras et al.'s EDM framework in continuous time [2509.07700]. The forward stochastic differential equation is
$$
dx = \mu(x,t)\,dt + \sigma(t)\,dW,
$$
with \(\mu(x,t)=0\), \(\sigma(t)=\sqrt{2}\,t\), and \(W\) a Wiener process. At \(t=0\), the process begins at the data distribution \(p_0(x)\); at \(t=T\), the distribution \(p_T(x)\approx \mathcal N(0,I)\).

The reverse-time SDE is
$$
dx = \bigl[\mu(x,t) - \sigma(t)^2 \nabla_x \log p_t(x)\bigr]\,dt + \sigma(t)\,d\overline W,
$$
and the corresponding probability-flow ODE is
$$
dx = \bigl[\mu(x,t) - \tfrac12 \sigma(t)^2 \nabla_x \log p_t(x)\bigr]\,dt.
$$
The score \(\nabla_x \log p_t(x)\) is approximated by a neural network \(s_\phi(x,t)\), reparameterised through a network \(v_\theta\) as
$$
s_\phi(x,t) = c_{\rm skip}(t)\,x \;+\; c_{\rm out}(t)\;v_\theta\bigl(c_{\rm in}(t)\,x,\;c_{\rm noise}(\sigma(t))\bigr),
$$
with
$$
c_{\rm skip}(t)=\frac{\sigma_{\rm data}^2}{\sigma_{\rm data}^2 + t^2}, \qquad
c_{\rm in}(t)=\frac{t}{\sqrt{\sigma_{\rm data}^2 + t^2}}, \qquad
c_{\rm out}(t)=\frac{1}{\sqrt{\sigma_{\rm data}^2 + t^2}}, \qquad
c_{\rm noise}(t)=\tfrac{\kappa}{4}\ln t,
$$
where \(\sigma_{\rm data}=0.5\) and \(\kappa=10^4\).

The EDM training loss is the weighted mean-squared error
$$
\mathcal L_{\rm EDM} =
\mathbb{E}_{x_0\sim p_0,\;t\sim\mathrm{noiseSchedule}}
\bigl[\lambda(t)\,\|s_\phi(x_t,t)-x_0\|_2^2\bigr],
$$
with
$$
\lambda(t)=\frac{t^2+\sigma_{\rm data}^2}{t\,\sigma_{\rm data}}.
$$

To reduce sampling cost, the work applies consistency distillation in the sense of Song et al. to collapse \(32\) EDM sampling steps into \(1\). With discretised times \(\epsilon=t_0<\cdots<t_N=T\), a student \(f_\zeta\) is trained against a fixed pre-trained teacher \(s_\phi\) through the self-consistency objective
$$
\mathcal L_{\rm CD}
=
\mathbb{E}\Bigl[
\|\,f_\zeta(x_{t_{n+1}},t_{n+1}) - f_{\zeta^-}(x_{t_n},t_n)\|_2^2
\Bigr],
$$
where \(\zeta^-\) is an exponential-moving-average of \(\zeta\). Sampling then requires only one call to \(f_{\zeta^-}\) [2509.07700].

## 4. Transformer architecture and implementation choices

The model architecture is based on DiT but is described as entirely end-to-end, with no VAE [2509.07700]. The input shower of shape \(9\times16\times45\) is split into non-overlapping \(3\times2\times3\) patches, producing
$$
(9/3)\cdot(16/2)\cdot(45/3)=360
$$
tokens. A shared Conv3D or linear layer projects each patch into an embedding dimension \(D=144\), and 3D sinusoidal positional embeddings are added by splitting the embedding space equally among the \(r\), \(\phi\), and \(z\) axes.

Condition injection is global rather than token-local. The diffusion time \(t\) and the physics conditions \(\hat H=[\hat E,\hat\theta,\sin\phi,\cos\phi,\mathrm{one\text{-}hot}\;G]\) are passed through small MLPs, concatenated into a condition vector \(c\), and supplied to every transformer block through adaptive LayerNorm in the adaLN-zero style. This makes the generative backbone conditional on both diffusion state and detector/kinematic metadata.

Each CaloDiT-2 block contains multi-head self-attention with \(8\) heads and head dimension \(18\), a SwiGLU MLP with hidden size \(4D=576\), and adaLN-zero with factorised affine conditioning on positional embedding and \(c\). Fixed scaling coefficients \(\alpha_{1,2}\) are replaced by \(\sigma(\alpha_i)\) gating for stability. The full model stacks \(4\) such blocks and has approximately \(2\) million parameters. Reported throughput for the diffusion variants is \(32\)-step EDM sampling at approximately \(170\) ms on GPU and \(1\)-step consistency-distilled sampling at approximately \(3\) ms [2509.07700].

The architectural profile is therefore modest in parameter count relative to contemporary transformer backbones, while the representation and conditioning pathway are tailored to sparse, structured calorimeter energy deposition data rather than to natural-image latents.

## 5. Multi-detector pre-training and adaptation regime

The pre-training strategy uses the LEMURS dataset with \(1\) million showers each of Par04-SiW, Par04-SciPb, ODD, and FCCeeCLD, split into \(900\,\mathrm{k}\) train and \(100\,\mathrm{k}\) validation events per detector [2509.07700]. EDM and consistency-distilled components are trained jointly in pre-training.

Adaptation to a novel detector is then demonstrated on FCCeeALLEGRO. The reported procedure is two-stage. First, the EDM model is fine-tuned on as few as \(1\,\mathrm{k}\) new showers for \(\lesssim 5\,\mathrm{k}\) steps at learning rate \(10^{-4}\). Second, the consistency-distilled model is trained using the fine-tuned EDM as teacher, with student initialisation from the pre-trained CD model.

The central quantitative claim is comparative efficiency: to reach the same Fréchet Physics Distance as training from scratch on \(25\,\mathrm{k}\) events for \(50\,\mathrm{k}\) steps, adaptation uses only \(1\,\mathrm{k}\) events and \(5\,\mathrm{k}\) steps, corresponding to \(25\times\) less data and \(10\times\) fewer steps, with EDM+CD combined yielding \(\lesssim 20\times\) less wall time [2509.07700]. The abstract additionally summarises the adaptation outcome as requiring up to \(25\times\) less data and \(20\times\) less training time.

A common misconception would be to interpret CaloDiT-2 as merely another detector-specific FastSim surrogate. The reported method explicitly combines detector-specific applicability with pre-training on multiple geometries and rapid transfer to previously unseen ones. The paper states, to the best of its knowledge, that this is the first pre-trained model to be published that allows adaptation in the context of particle shower simulations [2509.07700].

## 6. Reported performance, tradeoffs, and deployment in Geant4

On the single-detector Par04-SiW setting, the reported shower observables—longitudinal and transverse profiles, first and second moments, cell energy, and total energy—agree with Geant4 within statistical uncertainties for both the \(32\)-step EDM model and the \(1\)-step consistency-distilled model [2509.07700]. The stated caveat is that total energy shows slight under-coverage in the consistency-distilled model, although it remains acceptable. This is an important qualification because it distinguishes the faster sampler from the higher-fidelity teacher.

On Dataset-2 from the community-hosted CaloChallenge at \(500\) GeV for Par04-SiW, the reported metrics are as follows. For CaloDiT-2 EDM: \( \mathrm{AUC}_{\rm low}=0.594\pm0.002\), \( \mathrm{AUC}_{\rm high}=0.560\pm0.005\), \( \mathrm{FPD}\times10^3=20.1\pm0.7\), and \( \mathrm{KPD}\times10^3=0.09\pm0.09\). For CaloDiT-2 CD: \( \mathrm{AUC}_{\rm low}=0.598\pm0.002\), \( \mathrm{AUC}_{\rm high}=0.695\pm0.004\), \( \mathrm{FPD}\times10^3=68.8\pm3.4\), and \( \mathrm{KPD}\times10^3=0.39\pm0.15\). Relative to published models including CaloDREAM, CaloDiffusion, CaloINN, Calo-VQ, and CaloDiT-1, CaloDiT-2 is reported to lie at the Pareto frontier of accuracy versus speed.

For inference speed with batch size \(1\), the reported timings are \(6349\) ms on CPU and \(172\) ms on GPU for EDM, \(101\) ms on CPU and \(3\) ms on GPU for CD, and \(17323\) ms on CPU and \(640\) ms on GPU for CaloDiT-1 using DDPM. The consistency-distilled model is therefore approximately \(60\times\) faster than EDM on CPU and approximately \(200\times\) faster on GPU. Relative to Geant4 at approximately \(10\) s per shower, the CD model gives at least \(100\times\) speedup on GPU and at least \(100\times\) speedup even on a single-core CPU [2509.07700].

Deployment is part of the reported contribution. Pre-trained EDM and CD models are released in PyTorch, ONNX, and TorchScript, tagged V1 in the Geant4 GitLab. An extended Par04 example in Geant4 \(11.4.0\)-beta demonstrates three operations: generating virtual-mesh training data from any calorimeter, adapting the ONNX CD model to C++ via TorchScript or ONNXRuntime, and calling the model at each Geant4 shower step in place of the slow Geant4 physics process. Ready-to-use scripts automate adaptation, distillation, and export. The paper also states that the model is included in the Geant4 toolkit [2509.07700].

Source: https://www.emergentmind.com/topics/calodit-2