LEMURS Dataset: Multi-Detector EM Showers
- LEMURS dataset is a large-scale open repository of simulated electromagnetic calorimeter showers across five detector geometries using a universal grid representation.
- It employs Geant4, DD4hep, and custom ddfastsim libraries to generate ~5 million single-photon events with uniform energy and angular distributions.
- The dataset supports fast simulation and generative model benchmarking, reducing training data and time for cross-detector machine learning applications.
LEMURS, short for Large-scale multi-detector ElectroMagnetic Universal Representation of Showers, is an open dataset of simulated electromagnetic calorimeter showers for high-energy-physics fast-simulation research. It was designed to support the development and benchmarking of machine-learning-based simulators, including foundation-model approaches, by combining substantially greater statistics than earlier public benchmarks, a wider range of incident angles, and multiple detector geometries within a common detector-agnostic representation (McKeown et al., 5 Sep 2025). Produced with Geant4 and organized around a standardized Universal Grid Representation, LEMURS contains approximately single-photon showers across five ECAL configurations and has already been used as the pre-training substrate for CaloDiT-2, which is distributed in Geant4 v11.4.beta (McKeown et al., 5 Sep 2025).
1. Scope and detector coverage
LEMURS was constructed to replace narrow, single-geometry calorimeter benchmarks with a multi-detector corpus suitable for cross-geometry learning and transfer. The dataset covers five detector configurations, spanning both demonstrator calorimeters and more realistic collider detectors (McKeown et al., 5 Sep 2025).
| Detector | Geometry summary | Incident-energy range |
|---|---|---|
| Par04_SiW | Par04 demonstrator with alternating W–Si layers | to |
| Par04_SciPb | Par04 demonstrator with alternating Pb–scintillator layers | to |
| ODD | CERN Open Data Detector ECAL, 16-sided polygon, 48 tungsten/silicon sampling layers | to |
| FCC-ee CLD | 12-sided polygonal, 40-layer W–Si ECAL of the CLD concept | to |
| FCC-ee ALLEGRO | Noble-liquid, inclined-module ECAL with 12 readout layers and 1 500 cells per layer | to 0 |
All showers are generated for single photons entering the ECAL barrel from 1. The angular coverage is continuous rather than discrete: the polar angle 2 is sampled uniformly in 3 rad, corresponding to pseudorapidity 4, and the azimuth 5 is flat in 6 (McKeown et al., 5 Sep 2025).
A common misunderstanding is to treat LEMURS as merely a larger version of CaloChallenge dataset 2. The overlap is only partial. LEMURS retains the same Universal Grid Representation granularity for compatibility, but extends the benchmark from a single Par04_SiW geometry to five geometries, broadens the phase space to continuous angular conditioning, and raises the energy reach to 7 for the hadron-collider-style detectors (McKeown et al., 5 Sep 2025).
2. Simulation framework and event generation
The dataset was generated with Geant4 version 11.4.beta, invoked through the DD4hep/DDG4 simulation layer in the key4hep stack. A custom ddfastsim library implements the same Universal Grid Representation voxelization used in the Geant4 Par04 example, but extends it to multiple realistic detector geometries (McKeown et al., 5 Sep 2025).
The steering configuration uses the standard FTFP_BERT physics list for electromagnetic and hadronic processes. The paper specifies that photoelectric effect, Compton scattering, pair production, and bremsstrahlung are simulated at full Geant4 fidelity. Production was run on an HTCondor cluster, and hits in active volumes are scored onto a local cylindrical grid before export to HDF5 (McKeown et al., 5 Sep 2025).
The generation policy is deliberately uniform in the conditioning variables. For Par04_SiW, Par04_SciPb, and ODD, the GPSflat macro samples energies uniformly in 8; for FCC-ee CLD and ALLEGRO, it samples uniformly in 9. The absence of fixed angular slices is central to the dataset’s intended use in angle-conditioned generative modeling and transfer across detector designs (McKeown et al., 5 Sep 2025).
This suggests that LEMURS is not only a data resource but also an attempt to standardize the conditioning space of calorimeter simulation. By fixing the representation and broadening the detector family, it becomes possible to study whether a model learns shower physics that transfers across geometries rather than overfitting one detector-specific discretization.
3. Universal Grid Representation and file schema
A defining property of LEMURS is that showers are not stored in native detector readout coordinates. Instead, each event is expressed in a local cylindrical coordinate system whose 0-axis is aligned with the incident particle direction and whose origin is the ECAL entrance point 1. Transverse coordinates are 2 (McKeown et al., 5 Sep 2025).
Within this frame, the Universal Grid Representation (UGR) partitions the shower into a cylindrical voxel grid of fixed shape: 3
The corresponding HDF5 dataset showers has shape 4, where 5 is the number of showers in the file. The associated conditioning labels are stored as incident_energy, incident_theta, and incident_phi, each of shape 6. All datasets are float32 and gzip-9 compressed (McKeown et al., 5 Sep 2025).
The voxel energy is defined as
7
Longitudinal and radial coordinates can be expressed in detector-normalized form as
8
The voxel dimensions are chosen so that 9 and 0 in each detector; concrete examples given in the paper are 1, 2 for Par04_SiW and 3, 4 for ALLEGRO (McKeown et al., 5 Sep 2025).
This representation is detector-agnostic only in a controlled sense. The grid shape is standardized, but the physical calibration is not. The dataset explicitly leaves energy uncalibrated, with no sampling-fraction scaling, so downstream users may impose their own calibration conventions (McKeown et al., 5 Sep 2025). A plausible implication is that generative models trained on LEMURS can be compared at the representation level without forcing a single reconstruction prescription.
4. Statistics, phase-space coverage, and benchmark splits
For each of the five detectors, LEMURS provides approximately 5 showers, giving a total of approximately 6 showers. The paper notes that small variations arise from early conversions outside the ECAL, which are filtered out (McKeown et al., 5 Sep 2025).
The dataset also includes a fixed testing set for physics validation. For each detector, there are 1 000 showers at each of 8 discrete phase-space points, defined by:
- 7 for Par04/ODD or 8 for FCC-ee,
- 9,
- 0 (McKeown et al., 5 Sep 2025).
The dataset paper also reports sampling fractions—mean deposited energy divided by incident energy—ranging from approximately 3% for Par04 to approximately 14% for ALLEGRO (McKeown et al., 5 Sep 2025). These values summarize a substantive detector-dependent variation that is preserved despite the common voxelization.
Relative to CaloChallenge dataset 2, LEMURS changes both scale and task structure. CaloChallenge dataset 2 comprised approximately 300 k EM showers for only the Par04_SiW demonstrator and used fixed 1. LEMURS multiplies the statistics by a factor 2, expands to the full barrel 3 range and full 4, and adds four additional detector geometries (McKeown et al., 5 Sep 2025).
The CaloDiT-2 study describes the same detector collection with an experimental split of 900 k for training and 100 k for validation per detector (Raikwar et al., 9 Sep 2025). This is best read as a model-training protocol layered on top of the underlying dataset rather than as a conflicting alternative dataset definition.
5. Role in generative modeling and CaloDiT-2
LEMURS was designed explicitly for fast simulation and generative modeling. The dataset paper lists foundation-model pre-training, benchmarking of VAE, GAN, normalizing-flow, and diffusion architectures, studies of angle-conditioned simulation, and domain adaptation across detectors as intended applications (McKeown et al., 5 Sep 2025).
Its first major downstream use is CaloDiT-2, a diffusion model with transformer blocks that was pre-trained simultaneously on all five LEMURS detectors. In that setup, the model is conditioned on 5 together with a one-hot detector identifier, learning what the authors describe as a universal latent representation (McKeown et al., 5 Sep 2025). The companion CaloDiT-2 paper further states that this pre-training and adaptation strategy reduces the effort required to develop accurate models for new detectors, requiring up to 25x less data and 20x less training time (Raikwar et al., 9 Sep 2025).
For evaluation, the CaloDiT-2 workflow uses both low-level and physics-motivated observables: voxel-energy histograms, total deposited energy, longitudinal profile
6
transverse profile
7
and azimuthal profile
8
Quantitative metrics include AUC_low-level, AUC_high-level, FPD, KPD, and Precision & Recall / Density & Coverage (Raikwar et al., 9 Sep 2025).
The Geant4 integration is a further distinguishing feature. The dataset paper states that CaloDiT-2 is distributed with Geant4 v11.4.beta via a fast-simulation plugin, ddfastsim, and that replacing conventional Geant4 EM stepping with a call to CaloDiT-2 yields a speed-up of 9 for EM showers with sub-percent-level agreement on physics observables (McKeown et al., 5 Sep 2025).
6. Access, reproducibility, and terminological scope
LEMURS is distributed openly. The dataset paper states that all data and code for generation and analysis are openly accessible, and specifically points to the ddfastsim GitHub repository for generation, conversion, and plotting code (McKeown et al., 5 Sep 2025). The CaloDiT-2 paper additionally describes a public Zenodo release at DOI 10.5281/zenodo.17045562, with typical per-detector organization under /LEMURS/{detector}/train/ and /LEMURS/{detector}/val/, using NumPy .npy or HDF5 serialization plus JSON/YAML metadata (Raikwar et al., 9 Sep 2025).
Best-practice guidance in the dataset description emphasizes loading the showers array with shape 0, conditioning on 1 and optionally detector type, exploiting the strong sparsity of the voxel tensor, and using the provided fixed-point testing files for validation plots such as longitudinal and transverse profiles, energy-ratio distributions, and sparsity histograms (McKeown et al., 5 Sep 2025). The CaloDiT-2 study adds a model-specific preprocessing pipeline that clips voxel energies below 15.15 keV, applies a log transform and normalization, and rescales the conditioning variables as 2, 3, and 4 (Raikwar et al., 9 Sep 2025). Those preprocessing steps belong to a particular generative-model workflow rather than to the raw dataset specification.
The name itself invites potential confusion. The acronym LEMUR/LEMURS is also used for unrelated resources in other domains, including a primate face-recognition dataset (Deb et al., 2018), a behavioral-health study with differentially private synthetic releases (Ghasemizade et al., 30 Jun 2025), a multilingual legal-retrieval corpus (Ahmadi et al., 10 Feb 2026), and neural-network architecture datasets for AutoML (Goodarzi et al., 14 Apr 2025, Uzun et al., 7 Jul 2026). In high-energy physics, however, LEMURS refers specifically to the multi-detector electromagnetic-shower dataset based on the Universal Grid Representation (McKeown et al., 5 Sep 2025).