---
title: 'LEMURS Dataset: Multi-Detector EM Showers'
url: https://www.emergentmind.com/topics/lemurs-dataset
type: topic
---

# LEMURS Dataset: Multi-Detector EM Showers

LEMURS, short for **Large-scale multi-detector ElectroMagnetic Universal Representation of Showers**, is an open dataset of simulated electromagnetic calorimeter showers for high-energy-physics fast-simulation research. It was designed to support the development and benchmarking of machine-learning-based simulators, including foundation-model approaches, by combining substantially greater statistics than earlier public benchmarks, a wider range of incident angles, and multiple detector geometries within a common detector-agnostic representation [2509.05108]. Produced with Geant4 and organized around a standardized Universal Grid Representation, LEMURS contains approximately \(5\times 10^6\) single-photon showers across five ECAL configurations and has already been used as the pre-training substrate for CaloDiT-2, which is distributed in Geant4 v11.4.beta [2509.05108].

## 1. Scope and detector coverage

LEMURS was constructed to replace narrow, single-geometry calorimeter benchmarks with a multi-detector corpus suitable for cross-geometry learning and transfer. The dataset covers five detector configurations, spanning both demonstrator calorimeters and more realistic collider detectors [2509.05108].

| Detector | Geometry summary | Incident-energy range |
|---|---|---|
| Par04_SiW | Par04 demonstrator with alternating W–Si layers | \(1\,\mathrm{GeV}\) to \(1\,\mathrm{TeV}\) |
| Par04_SciPb | Par04 demonstrator with alternating Pb–scintillator layers | \(1\,\mathrm{GeV}\) to \(1\,\mathrm{TeV}\) |
| ODD | CERN Open Data Detector ECAL, 16-sided polygon, 48 tungsten/silicon sampling layers | \(1\,\mathrm{GeV}\) to \(1\,\mathrm{TeV}\) |
| FCC-ee CLD | 12-sided polygonal, 40-layer W–Si ECAL of the CLD concept | \(1\,\mathrm{GeV}\) to \(100\,\mathrm{GeV}\) |
| FCC-ee ALLEGRO | Noble-liquid, inclined-module ECAL with 12 readout layers and 1 500 cells per layer | \(1\,\mathrm{GeV}\) to \(100\,\mathrm{GeV}\) |

All showers are generated for **single photons** entering the ECAL barrel from \((0,0,0)\). The angular coverage is continuous rather than discrete: the polar angle \(\theta\) is sampled uniformly in \([0.87,\,2.27]\) rad, corresponding to pseudorapidity \(\eta\in[-0.76,0.76]\), and the azimuth \(\phi\) is flat in \([-\pi,\pi]\) [2509.05108].

A common misunderstanding is to treat LEMURS as merely a larger version of CaloChallenge dataset 2. The overlap is only partial. LEMURS retains the same Universal Grid Representation granularity for compatibility, but extends the benchmark from a single Par04_SiW geometry to five geometries, broadens the phase space to continuous angular conditioning, and raises the energy reach to \(1\,\mathrm{TeV}\) for the hadron-collider-style detectors [2509.05108].

## 2. Simulation framework and event generation

The dataset was generated with **Geant4 version 11.4.beta**, invoked through the **DD4hep/DDG4** simulation layer in the **key4hep** stack. A custom **ddfastsim** library implements the same Universal Grid Representation voxelization used in the Geant4 Par04 example, but extends it to multiple realistic detector geometries [2509.05108].

The steering configuration uses the standard **FTFP_BERT** physics list for electromagnetic and hadronic processes. The paper specifies that photoelectric effect, Compton scattering, pair production, and bremsstrahlung are simulated at full Geant4 fidelity. Production was run on an **HTCondor** cluster, and hits in active volumes are scored onto a local cylindrical grid before export to HDF5 [2509.05108].

The generation policy is deliberately uniform in the conditioning variables. For Par04_SiW, Par04_SciPb, and ODD, the **GPSflat** macro samples energies uniformly in \([1\,\mathrm{GeV},\,1\,\mathrm{TeV}]\); for FCC-ee CLD and ALLEGRO, it samples uniformly in \([1\,\mathrm{GeV},\,100\,\mathrm{GeV}]\). The absence of fixed angular slices is central to the dataset’s intended use in angle-conditioned generative modeling and transfer across detector designs [2509.05108].

This suggests that LEMURS is not only a data resource but also an attempt to standardize the *conditioning space* of calorimeter simulation. By fixing the representation and broadening the detector family, it becomes possible to study whether a model learns shower physics that transfers across geometries rather than overfitting one detector-specific discretization.

## 3. Universal Grid Representation and file schema

A defining property of LEMURS is that showers are not stored in native detector readout coordinates. Instead, each event is expressed in a **local cylindrical coordinate system** whose \(z\)-axis is aligned with the incident particle direction and whose origin is the ECAL entrance point \((x_0,y_0)\). Transverse coordinates are \((r,\phi)\) [2509.05108].

Within this frame, the **Universal Grid Representation (UGR)** partitions the shower into a cylindrical voxel grid of fixed shape:
\[
N\times R\times P = 45\times 9\times 16 = 6480.
\]
The corresponding HDF5 dataset `showers` has shape \((S,N,R,P)\), where \(S\) is the number of showers in the file. The associated conditioning labels are stored as `incident_energy`, `incident_theta`, and `incident_phi`, each of shape \((S,)\). All datasets are `float32` and `gzip-9` compressed [2509.05108].

The voxel energy is defined as
\[
E_{\mathrm{voxel}}(i,j,k)=\sum_{\mathrm{hits}\in \mathrm{voxel}(i,j,k)} E_{\mathrm{dep}}.
\]
Longitudinal and radial coordinates can be expressed in detector-normalized form as
\[
z' = z/X_0,\qquad r' = r/R_M.
\]
The voxel dimensions are chosen so that \(\Delta z \simeq 1.0\,X_0\) and \(\Delta r \simeq 1.0\,R_M\) in each detector; concrete examples given in the paper are \(\Delta r=4.65\,\mathrm{mm}\), \(\Delta z=3.4\,\mathrm{mm}\) for Par04_SiW and \(\Delta r=5\,\mathrm{mm}\), \(\Delta z=11\,\mathrm{mm}\) for ALLEGRO [2509.05108].

This representation is detector-agnostic only in a controlled sense. The grid shape is standardized, but the physical calibration is not. The dataset explicitly leaves energy **uncalibrated**, with no sampling-fraction scaling, so downstream users may impose their own calibration conventions [2509.05108]. A plausible implication is that generative models trained on LEMURS can be compared at the representation level without forcing a single reconstruction prescription.

## 4. Statistics, phase-space coverage, and benchmark splits

For each of the five detectors, LEMURS provides approximately **\(1\,000\,000\)** showers, giving a total of approximately **\(5\times 10^6\)** showers. The paper notes that small variations arise from early conversions outside the ECAL, which are filtered out [2509.05108].

The dataset also includes a fixed testing set for physics validation. For each detector, there are **1 000 showers at each of 8 discrete phase-space points**, defined by:
- \(E\in\{50,500\}\,\mathrm{GeV}\) for Par04/ODD or \(E\in\{5,50\}\,\mathrm{GeV}\) for FCC-ee,
- \(\theta\in\{1.57,2.10\}\,\mathrm{rad}\),
- \(\phi\in\{0.00,0.20\}\,\mathrm{rad}\) [2509.05108].

The dataset paper also reports **sampling fractions**—mean deposited energy divided by incident energy—ranging from approximately **3%** for Par04 to approximately **14%** for ALLEGRO [2509.05108]. These values summarize a substantive detector-dependent variation that is preserved despite the common voxelization.

Relative to **CaloChallenge dataset 2**, LEMURS changes both scale and task structure. CaloChallenge dataset 2 comprised approximately **300 k EM showers** for only the **Par04_SiW** demonstrator and used fixed \(\theta,\phi\). LEMURS multiplies the statistics by a factor \(\gtrsim 15\), expands to the full barrel \(\eta\) range and full \(\phi\), and adds four additional detector geometries [2509.05108].

The CaloDiT-2 study describes the same detector collection with an experimental split of **900 k for training** and **100 k for validation** per detector [2509.07700]. This is best read as a model-training protocol layered on top of the underlying dataset rather than as a conflicting alternative dataset definition.

## 5. Role in generative modeling and CaloDiT-2

LEMURS was designed explicitly for fast simulation and generative modeling. The dataset paper lists foundation-model pre-training, benchmarking of VAE, GAN, normalizing-flow, and diffusion architectures, studies of angle-conditioned simulation, and domain adaptation across detectors as intended applications [2509.05108].

Its first major downstream use is **CaloDiT-2**, a diffusion model with transformer blocks that was pre-trained simultaneously on all five LEMURS detectors. In that setup, the model is conditioned on \(E,\theta,\phi\) together with a one-hot detector identifier, learning what the authors describe as a universal latent representation [2509.05108]. The companion CaloDiT-2 paper further states that this pre-training and adaptation strategy reduces the effort required to develop accurate models for new detectors, requiring **up to 25x less data** and **20x less training time** [2509.07700].

For evaluation, the CaloDiT-2 workflow uses both low-level and physics-motivated observables: voxel-energy histograms, total deposited energy, longitudinal profile
\[
L(z)=\sum_{i_r,i_\phi} x_{i_r,i_\phi,i_z},
\]
transverse profile
\[
R(r)=\sum_{i_\phi,i_z} x_{i_r,i_\phi,i_z},
\]
and azimuthal profile
\[
\Phi(\phi)=\sum_{i_r,i_z} x_{i_r,i_\phi,i_z}.
\]
Quantitative metrics include **AUC_low-level**, **AUC_high-level**, **FPD**, **KPD**, and **Precision & Recall / Density & Coverage** [2509.07700].

The Geant4 integration is a further distinguishing feature. The dataset paper states that CaloDiT-2 is distributed with Geant4 v11.4.beta via a fast-simulation plugin, **ddfastsim**, and that replacing conventional Geant4 EM stepping with a call to CaloDiT-2 yields a speed-up of \(O(10^2\!-\!10^3)\) for EM showers with sub-percent-level agreement on physics observables [2509.05108].

## 6. Access, reproducibility, and terminological scope

LEMURS is distributed openly. The dataset paper states that all data and code for generation and analysis are openly accessible, and specifically points to the **ddfastsim GitHub repository** for generation, conversion, and plotting code [2509.05108]. The CaloDiT-2 paper additionally describes a public **Zenodo** release at DOI **10.5281/zenodo.17045562**, with typical per-detector organization under `/LEMURS/{detector}/train/` and `/LEMURS/{detector}/val/`, using NumPy `.npy` or HDF5 serialization plus JSON/YAML metadata [2509.07700].

Best-practice guidance in the dataset description emphasizes loading the `showers` array with shape \((\text{batch},45,9,16)\), conditioning on \((E,\theta,\phi)\) and optionally detector type, exploiting the strong sparsity of the voxel tensor, and using the provided fixed-point testing files for validation plots such as longitudinal and transverse profiles, energy-ratio distributions, and sparsity histograms [2509.05108]. The CaloDiT-2 study adds a model-specific preprocessing pipeline that clips voxel energies below **15.15 keV**, applies a log transform and normalization, and rescales the conditioning variables as \(\hat E=E/E_{\max}\), \(\hat\theta=\theta/\pi\), and \(\hat\phi=(\sin\phi,\cos\phi)\) [2509.07700]. Those preprocessing steps belong to a particular generative-model workflow rather than to the raw dataset specification.

The name itself invites potential confusion. The acronym **LEMUR/LEMURS** is also used for unrelated resources in other domains, including a primate face-recognition dataset [1804.08790], a behavioral-health study with differentially private synthetic releases [2507.02971], a multilingual legal-retrieval corpus [2602.09570], and neural-network architecture datasets for AutoML [2504.10552; 2607.06839]. In high-energy physics, however, **LEMURS** refers specifically to the multi-detector electromagnetic-shower dataset based on the Universal Grid Representation [2509.05108].

Source: https://www.emergentmind.com/topics/lemurs-dataset