Higgsformer: Reconstruction-Free LHC Event Classification
- Higgsformer is a set-based transformer that classifies LHC Higgs events using only raw inner tracker hit coordinates, eliminating the need for full reconstruction.
- It employs multi-head self-attention and geometric augmentations to achieve competitive ROC AUC scores in discriminating t¯tH events from the dominant t¯t background.
- The method enables fast inference (<10ms per event) and presents efficient deployment potential for online trigger strategies in high-energy physics.
Higgsformer is a lightweight, set-based transformer for event-level classification directly from raw inner tracker hits, without reconstructing tracks, jets, or other physics objects. It was introduced in the context of Higgs event classification at the Large Hadron Collider as a reconstruction-free alternative to conventional analysis chains, with the benchmark task of separating events from the dominant background when . In that setting, the model operates solely on inner tracker hit coordinates and nevertheless achieves competitive discrimination, reaching ROC AUC at zero pileup with geometric augmentation (Caron et al., 26 Aug 2025).
1. Problem setting and motivation
The Higgsformer study targets one of the most challenging Higgs signatures at the LHC: distinguishing from in the channel. These final states share very similar topologies—two top quarks plus additional hadronic activity—with only subtle differences from the extra -jets produced by the Higgs decay. The central methodological claim is that this discrimination can be attempted directly from raw detector response, rather than from reconstructed high-level objects (Caron et al., 26 Aug 2025).
Traditional analyses in this regime rely on full reconstruction, including tracking, calorimetry, and -tagging. In the Higgsformer formulation, that dependence is treated as both a strength and a limitation: reconstruction provides strong inductive biases, but it may also discard low-level information present in the detector response. The reconstruction-free program is therefore motivated by three explicit considerations: reducing latency and computational cost by bypassing heavy reconstruction, preserving low-level geometric and timing details that may be informative for difficult classifications, and enabling trigger or online strategies that operate closer to the detector.
A common misconception is that Higgsformer should be evaluated as a drop-in replacement for full-detector, object-based classification. The reported comparison is narrower. Higgsformer uses only inner tracker hits, while the strongest baselines use reconstructed objects from the full detector, including calorimetry and muon information. The relevant significance of the result is therefore not superiority over mature object-based workflows, but the demonstration that nontrivial Higgs sensitivity can emerge directly from raw hit patterns.
2. Event generation, detector simulation, and hit representation
The data pipeline begins with event generation in \texttt{Pythia8}. Signal corresponds to with 0, and background corresponds to 1. Both 2 and 3 are included; ISR and FSR are enabled; parton-level 4; and hadron-level decays are fully enabled. To emulate the LHC beam spot, vertex smearing is applied with 5, 6, and 7, using fixed seed 42. Detector simulation is performed with ACTS/Fatras through an ACTS GenericDetector with TrackML-like geometry in a 8 magnetic field, including multiple scattering, energy loss, and interactions. Digitization uses realistic pixel and strip response via ACTS JSON configurations, producing digitized hits (Caron et al., 26 Aug 2025).
The raw-hit preprocessing stage converts local sensor coordinates into global 9, attaches TrackML-style metadata such as volume, layer, and module indices, removes malformed hits and zero-charge particles, computes per-track 0, and assigns per-hit weights inspired by TrackML. However, the classifier itself does not consume the detector tags. Higgsformer uses only the three spatial coordinates 1 per hit. Each event is represented as a variable-length set of inner tracker hits, with occupancy ranging from hundreds to thousands depending on event complexity and pileup. Padding and masking are used for batching.
The study evaluates pileup levels of 0, 5, and 20 additional interactions per bunch crossing. For each pileup level, 40,000 balanced events are produced, comprising 20,000 signal and 20,000 background events, and stored in ROOT and CSV. Dataset-size studies are performed at 10k, 20k, and 40k events, with 1k reserved for validation and 1k for testing in each case and the remainder used for training.
This representation is deliberately minimal. There is no heavy feature engineering, and no handcrafted kinematic abstraction mediates between the detector and the classifier. The study therefore isolates how much event-level information can be extracted from raw spatial hit structure alone.
3. Architecture and mathematical formulation
Higgsformer is adapted from a previously developed Trackformer designed for inner tracker hit-to-track assignment. In the Higgs classification setting, the same attention backbone is repurposed for event-level classification from raw hits rather than track construction. The architecture is explicitly set-based: hits are treated as an unordered collection, and permutation invariance is achieved through self-attention over hit tokens followed by pooled aggregation (Caron et al., 26 Aug 2025).
Each hit is embedded by a linear map from its 3D position 2 to a 3-dimensional token representation. Multi-head self-attention is then applied using
4
with exact FlashAttention used for speed and memory efficiency. The reported small configuration uses 5 transformer encoder layers, 6 attention heads, and hidden size 7, while larger variants were tested up to 8, 9, and 0. Standard encoder components are retained, including layer normalization and position-wise feed-forward blocks with nonlinearity; dropout is set to 0.3 and weight decay is applied.
No handcrafted positional encoding is added beyond the raw coordinates themselves. Consequently, detector geometry and symmetries must be captured implicitly through the learned embedding and attention mechanism rather than through engineered geometric priors.
Event-level aggregation is performed by masked mean pooling across hit embeddings. The pooled vector is mapped by a linear layer to a scalar logit 1, which is converted into a probability 2. Training uses binary cross-entropy with logits,
3
aggregated over events. To counter a background bias in the logits, the positive-class weight is set to 4.
4. Training protocol and baseline comparators
Optimization uses AdamW with learning rate 5 and batch size 64. Training runs for up to 500 epochs with early stopping of patience 100 epochs based on validation AUC. Mixed precision (AMP) and FlashAttention are used throughout. The regularization scheme consists of dropout 0.3 and weight decay. Two online geometric augmentations are applied only to training data: rotations in the transverse plane 6 through transformations of 7, and forward–backward reflections 8. Although the datasets are balanced, the BCE positive-class weight of 1.5 is retained to mitigate slight logit bias (Caron et al., 26 Aug 2025).
For comparison with reconstruction-based workflows, the same events are reconstructed with Delphes 3.5.0 using an ATLAS-like card. Jets are clustered with anti-9 and 0; standard detector resolutions and 1-tagging are used; and reconstructed objects are ordered by 2 and padded per category with masks. Two baseline classifier families are studied. The first consists of multilayer perceptrons trained on fixed-length arrays of object features. The second is the Particle Transformer (ParT), which processes variable-length sequences of reconstructed objects.
The comparison is informative but not fully matched in detector information content. The object-based baselines use the full detector—inner tracker, calorimetry, and muon systems—whereas Higgsformer uses only inner tracker hits. The resulting performance gap must therefore be interpreted as a comparison between reconstruction-free inner-tracker classification and full-detector object-level classification, not as a like-for-like architecture benchmark.
5. Empirical performance and comparative results
The primary metric is ROC AUC, with ROC curves defined by true positive rate 3 versus false positive rate 4 as the decision threshold is swept. The significance improvement characteristic may be defined as 5, but SIC is not reported numerically. Random guessing corresponds to AUC 6. Operating points are not explicitly tabulated, and error bars across random seeds are not provided.
Performance scales with dataset size at zero pileup. For Higgsformer-small without augmentation, the AUC rises from approximately 0.704 at 10k events to 0.757 at 20k and 0.779 at 40k. Geometric augmentation improves each setting, yielding approximately 0.721, 0.764, and 0.792, respectively (Caron et al., 26 Aug 2025).
| Setting | AUC | Condition |
|---|---|---|
| 10k events | 7 | PU = 0, no augmentation |
| 20k events | 8 | PU = 0, no augmentation |
| 40k events | 9 | PU = 0, no augmentation |
| 10k events | 0 | PU = 0, with augmentation |
| 20k events | 1 | PU = 0, with augmentation |
| 40k events | 2 | PU = 0, with augmentation |
Pileup degrades performance but does not eliminate discrimination. With 38k training events and Higgsformer-small without augmentation, AUC is approximately 0.779 at PU 3, 0.731 at PU 4, and 0.654 at PU 5. Even at PU 6, performance remains well above random.
A naive counts-only baseline using 7 per event reaches only approximately 0.621, 0.560, and 0.530 at PU 8, 5, and 20, respectively. Higgsformer exceeds this baseline by approximately 9 to 0 AUC depending on pileup, indicating that the network exploits structured spatial information rather than merely occupancy.
The object-based baselines remain stronger in absolute terms:
| Model | AUC | Inputs |
|---|---|---|
| MLP | 1–0.96 | Full detector objects |
| MLP | 2–0.86 | Only 3-jets |
| ParT | 4 | Full detector objects |
These results establish the central positioning of Higgsformer. Under matched physics generation, full-detector object-based methods still outperform. However, Higgsformer closes a significant fraction of the gap without any reconstruction and while using only inner tracker hits.
6. Interpretability, deployment relevance, and limitations
The study includes a leave-one-hit-out importance analysis in which individual hits are removed and the corresponding change in loss is measured. Hits identified as descendants of the Higgs decay through HepMC 5 ROOT 6 CSV matching have higher average importance than non-Higgs hits. With 38k augmented training events, the mean importance is approximately 7 versus 8, and in 72 of 98 events 9. With 8k augmented training events, the corresponding means are approximately 0.00222 and 0.00197, with 51 of 98 events satisfying the same inequality. The widening gap with larger training size indicates increasing sensitivity to Higgs-origin signal content.
A complementary saliency visualization finds that the top-10 most important hits per event form increasingly symmetric, cylindrical patterns as the training data scale. This suggests that the model learns detector geometry and exploits structured hit configurations consistent with 0-jet-rich topologies rather than relying on unstructured occupancy artifacts alone.
The practical deployment argument rests on latency and simplicity. Higgsformer-small runs in less than 10 ms per event on an NVIDIA H100 GPU at PU 1, compared with classical CPU-based tracking at approximately 1 s per event. The shallow depth and small hidden dimension keep memory demands low, and FlashAttention improves efficiency on variable-length hit sets. A plausible implication is that such models could serve as fast filtering stages or components of hybrid trigger pipelines that pre-select events before full reconstruction.
Several limitations are explicit. Only inner tracker hits are used; domain shifts from changed detector conditions or higher pileup may degrade performance; training sizes are modest, at most 38k training events; and performance decreases substantially at PU 2. The paper therefore identifies calibration under systematic variation, larger datasets, deeper models, explicit pileup modeling or denoising, extension to other Higgs channels such as 3 and 4, and hybrid combinations with partial reconstruction or learned tracklet formation as natural future directions. Reproducibility is supported by a public generation and hit-processing pipeline at https://github.com/EugeneShalli/hits-gen, together with detailed reporting of the event generation stack, detector simulation, data splits, and principal hyperparameters (Caron et al., 26 Aug 2025).