Papers
Topics
Authors
Recent
Search
2000 character limit reached

CaloDiT-2: Transformer Diffusion for Fast Calorimeter Simulation

Updated 10 July 2026
  • The paper introduces CaloDiT-2, a transformer-based diffusion model that matches Geant4 shower observables while dramatically accelerating simulation speed.
  • It leverages detector-agnostic representations and pre-training on multiple detector geometries to enable rapid adaptation to new calorimeter designs.
  • Consistency distillation reduces sampling steps from 32 to 1, achieving up to 200× faster inference on GPU and significantly less training data and time.

CaloDiT-2 is a transformer-based diffusion model for fast simulation of electromagnetic calorimeter showers in collider experiments, introduced in "A Generalisable Generative Model for Multi-Detector Calorimeter Simulation" (Raikwar et al., 9 Sep 2025). It is designed as a generative surrogate for Geant4 calorimeter simulation, with an emphasis not only on speed and fidelity but also on cross-detector generalisation. The model is formulated around detector-agnostic data representations, transformer diffusion blocks, and a pre-training-plus-adaptation regime on multiple detector geometries. In the reported experiments, it produces showers in approximately $100$ ms on CPU and approximately $3$ ms on GPU, matches Geant4 on critical shower observables, and adapts to novel detectors with substantially reduced data and training requirements (Raikwar et al., 9 Sep 2025).

1. Research setting and stated objective

High-energy physics experiments, including those at the Large Hadron Collider, rely on the Geant4 toolkit for high-fidelity calorimeter shower simulation, but the CPU cost is described as seconds per shower and becomes prohibitive at high luminosity (Raikwar et al., 9 Sep 2025). CaloDiT-2 is positioned as a response to this computational bottleneck: a diffusion model using transformer blocks that directly emulates detector responses while preserving the observables that matter for calorimeter simulation.

The reported objective is broader than accelerating a single fixed detector geometry. CaloDiT-2 can be applied to specific geometries, as is the case for other models explored for the task, but its stated strength lies in generalisation across detectors through pre-training on multiple geometries and rapid adaptation to new ones. The work frames this as an ambition toward a "foundation model" for FastSim: train once on varied geometries, then adapt cheaply to novel detectors under design (Raikwar et al., 9 Sep 2025).

The scope of the presented model is specifically electromagnetic calorimeter shower generation. The conditioning variables are defined for incident photons with energy E∈[1,1000]E \in [1,1000] GeV and angular variables ϕ∈[0,2π]\phi \in [0,2\pi] and θ∈[0.87,2.27]\theta \in [0.87,2.27]. A plausible implication is that the reported generalisation claims are anchored to this simulation regime rather than to calorimeter simulation in full generality.

2. Detector-agnostic representation and conditioning interface

A central component of CaloDiT-2 is a detector-agnostic data representation built from a virtual cylindrical scoring mesh aligned with the incident particle, exploiting approximate azimuthal symmetry (Raikwar et al., 9 Sep 2025). In the Par04/LEMURS setup, voxels are defined in (r,ϕ,z)(r,\phi,z) with fixed resolution 9×16×459\times16\times45. This mesh is independent of any real read-out geometry: all energy deposits from Geant4 are scored into the mesh, yielding a tensor x∈R9×16×45x \in \mathbb{R}^{9\times16\times45}.

This representation separates the learning problem from any one detector segmentation scheme. Instead of binding the network to a specific read-out layout, the model operates on a common volumetric coordinate system into which multiple detectors can be projected. That design is the basis for the multi-detector pre-training strategy.

The conditioning variables are the incident photon energy EE, the particle angles ϕ\phi and $3$0, and a detector identifier $3$1, where $3$2 is the number of pre-training detectors and the extra bit flags an "unknown" geometry. Voxel preprocessing is given by

$3$3

where $3$4 and $3$5 are the mean and standard deviation of log-energies over the entire pre-training set, and $3$6. The scalar conditions are normalised as $3$7 and $3$8, while $3$9 is encoded as E∈[1,1000]E \in [1,1000]0 to preserve cyclicity (Raikwar et al., 9 Sep 2025).

The use of an "unknown" geometry flag is notable because it makes the conditioning scheme explicitly compatible with adaptation scenarios. This suggests that the representation is intended not merely for interpolation among known detectors but also for transfer to unseen geometries.

3. Diffusion formulation and consistency distillation

CaloDiT-2 follows Karras et al.'s EDM framework in continuous time (Raikwar et al., 9 Sep 2025). The forward stochastic differential equation is

E∈[1,1000]E \in [1,1000]1

with E∈[1,1000]E \in [1,1000]2, E∈[1,1000]E \in [1,1000]3, and E∈[1,1000]E \in [1,1000]4 a Wiener process. At E∈[1,1000]E \in [1,1000]5, the process begins at the data distribution E∈[1,1000]E \in [1,1000]6; at E∈[1,1000]E \in [1,1000]7, the distribution E∈[1,1000]E \in [1,1000]8.

The reverse-time SDE is

E∈[1,1000]E \in [1,1000]9

and the corresponding probability-flow ODE is

ϕ∈[0,2π]\phi \in [0,2\pi]0

The score ϕ∈[0,2π]\phi \in [0,2\pi]1 is approximated by a neural network ϕ∈[0,2π]\phi \in [0,2\pi]2, reparameterised through a network ϕ∈[0,2π]\phi \in [0,2\pi]3 as

ϕ∈[0,2π]\phi \in [0,2\pi]4

with

ϕ∈[0,2π]\phi \in [0,2\pi]5

where ϕ∈[0,2π]\phi \in [0,2\pi]6 and ϕ∈[0,2π]\phi \in [0,2\pi]7.

The EDM training loss is the weighted mean-squared error

ϕ∈[0,2π]\phi \in [0,2\pi]8

with

ϕ∈[0,2π]\phi \in [0,2\pi]9

To reduce sampling cost, the work applies consistency distillation in the sense of Song et al. to collapse θ∈[0.87,2.27]\theta \in [0.87,2.27]0 EDM sampling steps into θ∈[0.87,2.27]\theta \in [0.87,2.27]1. With discretised times θ∈[0.87,2.27]\theta \in [0.87,2.27]2, a student θ∈[0.87,2.27]\theta \in [0.87,2.27]3 is trained against a fixed pre-trained teacher θ∈[0.87,2.27]\theta \in [0.87,2.27]4 through the self-consistency objective

θ∈[0.87,2.27]\theta \in [0.87,2.27]5

where θ∈[0.87,2.27]\theta \in [0.87,2.27]6 is an exponential-moving-average of θ∈[0.87,2.27]\theta \in [0.87,2.27]7. Sampling then requires only one call to θ∈[0.87,2.27]\theta \in [0.87,2.27]8 (Raikwar et al., 9 Sep 2025).

4. Transformer architecture and implementation choices

The model architecture is based on DiT but is described as entirely end-to-end, with no VAE (Raikwar et al., 9 Sep 2025). The input shower of shape θ∈[0.87,2.27]\theta \in [0.87,2.27]9 is split into non-overlapping (r,ϕ,z)(r,\phi,z)0 patches, producing

(r,ϕ,z)(r,\phi,z)1

tokens. A shared Conv3D or linear layer projects each patch into an embedding dimension (r,ϕ,z)(r,\phi,z)2, and 3D sinusoidal positional embeddings are added by splitting the embedding space equally among the (r,ϕ,z)(r,\phi,z)3, (r,ϕ,z)(r,\phi,z)4, and (r,ϕ,z)(r,\phi,z)5 axes.

Condition injection is global rather than token-local. The diffusion time (r,ϕ,z)(r,\phi,z)6 and the physics conditions (r,ϕ,z)(r,\phi,z)7 are passed through small MLPs, concatenated into a condition vector (r,ϕ,z)(r,\phi,z)8, and supplied to every transformer block through adaptive LayerNorm in the adaLN-zero style. This makes the generative backbone conditional on both diffusion state and detector/kinematic metadata.

Each CaloDiT-2 block contains multi-head self-attention with (r,ϕ,z)(r,\phi,z)9 heads and head dimension 9×16×459\times16\times450, a SwiGLU MLP with hidden size 9×16×459\times16\times451, and adaLN-zero with factorised affine conditioning on positional embedding and 9×16×459\times16\times452. Fixed scaling coefficients 9×16×459\times16\times453 are replaced by 9×16×459\times16\times454 gating for stability. The full model stacks 9×16×459\times16\times455 such blocks and has approximately 9×16×459\times16\times456 million parameters. Reported throughput for the diffusion variants is 9×16×459\times16\times457-step EDM sampling at approximately 9×16×459\times16\times458 ms on GPU and 9×16×459\times16\times459-step consistency-distilled sampling at approximately x∈R9×16×45x \in \mathbb{R}^{9\times16\times45}0 ms (Raikwar et al., 9 Sep 2025).

The architectural profile is therefore modest in parameter count relative to contemporary transformer backbones, while the representation and conditioning pathway are tailored to sparse, structured calorimeter energy deposition data rather than to natural-image latents.

5. Multi-detector pre-training and adaptation regime

The pre-training strategy uses the LEMURS dataset with x∈R9×16×45x \in \mathbb{R}^{9\times16\times45}1 million showers each of Par04-SiW, Par04-SciPb, ODD, and FCCeeCLD, split into x∈R9×16×45x \in \mathbb{R}^{9\times16\times45}2 train and x∈R9×16×45x \in \mathbb{R}^{9\times16\times45}3 validation events per detector (Raikwar et al., 9 Sep 2025). EDM and consistency-distilled components are trained jointly in pre-training.

Adaptation to a novel detector is then demonstrated on FCCeeALLEGRO. The reported procedure is two-stage. First, the EDM model is fine-tuned on as few as x∈R9×16×45x \in \mathbb{R}^{9\times16\times45}4 new showers for x∈R9×16×45x \in \mathbb{R}^{9\times16\times45}5 steps at learning rate x∈R9×16×45x \in \mathbb{R}^{9\times16\times45}6. Second, the consistency-distilled model is trained using the fine-tuned EDM as teacher, with student initialisation from the pre-trained CD model.

The central quantitative claim is comparative efficiency: to reach the same Fréchet Physics Distance as training from scratch on x∈R9×16×45x \in \mathbb{R}^{9\times16\times45}7 events for x∈R9×16×45x \in \mathbb{R}^{9\times16\times45}8 steps, adaptation uses only x∈R9×16×45x \in \mathbb{R}^{9\times16\times45}9 events and EE0 steps, corresponding to EE1 less data and EE2 fewer steps, with EDM+CD combined yielding EE3 less wall time (Raikwar et al., 9 Sep 2025). The abstract additionally summarises the adaptation outcome as requiring up to EE4 less data and EE5 less training time.

A common misconception would be to interpret CaloDiT-2 as merely another detector-specific FastSim surrogate. The reported method explicitly combines detector-specific applicability with pre-training on multiple geometries and rapid transfer to previously unseen ones. The paper states, to the best of its knowledge, that this is the first pre-trained model to be published that allows adaptation in the context of particle shower simulations (Raikwar et al., 9 Sep 2025).

6. Reported performance, tradeoffs, and deployment in Geant4

On the single-detector Par04-SiW setting, the reported shower observables—longitudinal and transverse profiles, first and second moments, cell energy, and total energy—agree with Geant4 within statistical uncertainties for both the EE6-step EDM model and the EE7-step consistency-distilled model (Raikwar et al., 9 Sep 2025). The stated caveat is that total energy shows slight under-coverage in the consistency-distilled model, although it remains acceptable. This is an important qualification because it distinguishes the faster sampler from the higher-fidelity teacher.

On Dataset-2 from the community-hosted CaloChallenge at EE8 GeV for Par04-SiW, the reported metrics are as follows. For CaloDiT-2 EDM: EE9, ϕ\phi0, ϕ\phi1, and ϕ\phi2. For CaloDiT-2 CD: ϕ\phi3, ϕ\phi4, ϕ\phi5, and ϕ\phi6. Relative to published models including CaloDREAM, CaloDiffusion, CaloINN, Calo-VQ, and CaloDiT-1, CaloDiT-2 is reported to lie at the Pareto frontier of accuracy versus speed.

For inference speed with batch size ϕ\phi7, the reported timings are ϕ\phi8 ms on CPU and ϕ\phi9 ms on GPU for EDM, $3$00 ms on CPU and $3$01 ms on GPU for CD, and $3$02 ms on CPU and $3$03 ms on GPU for CaloDiT-1 using DDPM. The consistency-distilled model is therefore approximately $3$04 faster than EDM on CPU and approximately $3$05 faster on GPU. Relative to Geant4 at approximately $3$06 s per shower, the CD model gives at least $3$07 speedup on GPU and at least $3$08 speedup even on a single-core CPU (Raikwar et al., 9 Sep 2025).

Deployment is part of the reported contribution. Pre-trained EDM and CD models are released in PyTorch, ONNX, and TorchScript, tagged V1 in the Geant4 GitLab. An extended Par04 example in Geant4 $3$09-beta demonstrates three operations: generating virtual-mesh training data from any calorimeter, adapting the ONNX CD model to C++ via TorchScript or ONNXRuntime, and calling the model at each Geant4 shower step in place of the slow Geant4 physics process. Ready-to-use scripts automate adaptation, distillation, and export. The paper also states that the model is included in the Geant4 toolkit (Raikwar et al., 9 Sep 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to CaloDiT-2.