---
title: Quark–Gluon Tagging in Collider Physics
url: https://www.emergentmind.com/topics/quark-gluon-tagging
type: topic
---

# Quark–Gluon Tagging in Collider Physics

Searching arXiv for relevant quark–gluon tagging papers to ground the article in published work.
Quark–gluon tagging is the problem of distinguishing jets initiated by quarks from jets initiated by gluons using the internal radiation pattern of the reconstructed jet. It is important in many collider analyses because signal processes are often quark-enriched while backgrounds are more gluon-enriched, so improved discrimination can enhance signal selection, background rejection, and overall analysis sensitivity [1409.3072]. Across the literature, the subject spans detector-level taggers used by CMS and ATLAS, analytically controlled observables based on perturbative QCD, unsupervised density-estimation methods, and constituent-level deep learning architectures, with persistent emphasis on calibration, topology dependence, and theoretical systematics [2308.00716].

## 1. Operational definition and scope

A common practical definition labels a jet by the hard parton that initiates the shower in the leading-order parton-shower picture [1106.3076]. At the same time, several studies stress that there is no single universal, gauge-invariant hadron-level definition of a “quark jet” or “gluon jet,” so the problem is inherently operational and process-dependent [2103.09103]. This ambiguity is one reason experimental calibrations rely heavily on quark-enriched and gluon-enriched event samples rather than on event-by-event truth labels in data [1409.3072].

The calibration issue is sharpened by topology dependence. Because quarks and gluons are colored while only colorless hadrons are measured, the radiation pattern inside a jet depends on the rest of the event, including color connections, nearby jets, underlying event, and ISR [1810.05653]. In simulation studies, same-flavor topology dependence is much smaller than quark-versus-gluon separation, and quark and gluon jets are approximately universal up to $\mathcal{O}(10\%)$ corrections, with values more like $\sim 2\%$ for IRC-safe observables in PYTHIA [1810.05653]. This does not remove the problem; it establishes the scale on which transferability between processes should be assessed.

Experimental practice therefore uses flavor-enriched regions. CMS used $Z+$jets as a quark-enriched control region and dijets as a gluon-enriched control region, without assigning direct event-by-event truth labels in data [1409.3072]. ATLAS used forward and central dijet subsamples, exploiting the fact that the more forward jet is quark-enriched while the more central jet is gluon-enriched at high momentum fraction $x$ [2308.00716]. A plausible implication is that quark–gluon tagging is best understood not as exact parton identification, but as calibrated discrimination between statistically different jet populations.

## 2. Physical origin of discrimination and the observable basis

The physical basis is standard QCD color-charge scaling. Since gluons carry a larger color charge than quarks, gluon-initiated jets tend to be wider, have higher multiplicity, and show softer or more uniform fragmentation, while quark jets are narrower, contain fewer constituents, and have harder leading fragments [1409.3072]. The frequently quoted color-factor ratio is $C_A/C_F = 9/4$, with $C_A=3$ and $C_F=4/3$ [2308.00716].

The observable basis developed in the literature reflects these differences. Early work separated discriminants into discrete observables, such as charged track multiplicity and subjet multiplicity, and continuous shape observables, such as jet mass, radial moments, angularities, pull, eccentricity, and planar flow [1106.3076]. Among the simplest and most widely reused variables are charged track multiplicity and girth,
$$
g = \sum_{i \in \mathrm{jet}} \frac{p_T^i}{p_T^{\text{jet}}}\,|r_i| ,
$$
which encode multiplicity and radial energy flow, respectively [1106.3076].

A unifying language is provided by generalized angularities,
$$
\lambda^\kappa_\beta = \sum_{i\in \text{jet}} z_i^\kappa \,\theta_i^\beta ,
$$
with benchmark cases
$$
(0,0)\to \text{multiplicity},\qquad
(2,0)\to p_T^D,\qquad
(1,0.5)\to \text{LHA},\qquad
(1,1)\to \text{width},\qquad
(1,2)\to \text{mass/thrust}
$$
[1704.03878]. This organization makes explicit the distinction between IRC-safe and IRC-unsafe observables and clarifies which measurements are expected to be most sensitive to hadronization and detector thresholds.

CMS built a likelihood discriminator from three particularly robust inputs: multiplicity, the jet energy-sharing variable $p_D$, and the angular spread measured by the minor axis $\sigma_2$ of the jet in the $\eta$–$\phi$ plane [1409.3072]. ATLAS later combined $\ntrk$, jet track width $\wtrk$, and the two-point energy correlation function $\cbeta$ with $\beta=0.2$ in a BDT, while also studying a pure track-multiplicity tagger [2308.00716]. Other studies used $w_\text{PF}$, $p_TD$, $C_{0.2}$, $N_{95}$, and $x_\text{max}$ as baseline high-level observables for detector-level multivariate taggers [1812.09223]. The convergence of these different programs is notable: multiplicity, width-like observables, and energy-sharing observables recur across analytic, experimental, and machine-learning treatments.

## 3. Experimental taggers and calibration strategies

The 8 TeV CMS likelihood tagger is a representative detector-level implementation. It restricted charged PF candidates to tracks compatible with the primary vertex and neutral PF candidates to those with $p_T>1$ GeV, then constructed the likelihood in bins of jet transverse momentum $p_T$ and pileup density $\rho$, separately for central jets with $|\eta|<2.4$ and forward jets with $2.4<|\eta|<4.7$ [1409.3072]. Validation used a $Z+$jets selection with $\Delta\phi(Z,\text{jet})>2.5\ \text{rad}$ for a quark-enriched sample and a back-to-back dijet selection for a gluon-enriched sample. Performance was studied with ROC curves, efficiencies after cuts such as likelihood discriminant $>0.5$, and comparisons of input-variable and discriminant distributions between data and 0.8.6 simulation or 0.8 simulation [1409.3072].

A central CMS result was that raw simulation did not perfectly reproduce data, so a shape-uncertainty treatment based on a smearing function was introduced:
$$
g(x,a,b)=\frac{1}{2}\tanh\left(a\,\mathrm{arctanh}(2x-1)+b\right) + \frac{1}{2}.
$$
Applied independently to quark and gluon distributions, with parameters obtained from a $\chi^2$ minimization, this transformed the simulated discriminant output while keeping it in $[0,1]$ and brought simulation into better agreement with data [1409.3072]. The same framework also reconciled the fact that data had worse discrimination than 0.8.6 simulation but better discrimination than 0.8 simulation.

ATLAS extended this program to 140 fb$^{-1}$ of $pp$ collisions at $\sqrt{s}=13$ TeV, focusing on jets with $500~\mathrm{GeV} \le p_T \le 2~\mathrm{TeV}$ [2308.00716]. Two taggers were studied: a single-variable $\ntrk$ tagger and a BDT trained on jet $p_T$, $\ntrk$, $\wtrk$, and $\cbeta$ with $\beta=0.2$, using LightGBM with Optuna tuning on $\sim 60$ million simulated two-jet events; the final model used 224 leaves after 100 boosting iterations [2308.00716]. Quark- and gluon-enriched subsamples were defined by jet pseudorapidity, and a matrix method extracted underlying quark and gluon score distributions in data from the forward and central mixtures.

ATLAS defined working points at fixed quark efficiency in nominal PYTHIA simulation, specifically 50%, 60%, 70%, and 80% [2308.00716]. At the 50% quark-efficiency working point, the $\ntrk$ tagger rejected about 90% of gluon jets and the BDT tagger rejected about 93% of gluon jets. Data-to-MC scale factors for both taggers lay in the range 0.92 to 1.02, with total uncertainty about 20%, growing at higher $p_T$; the main uncertainty was theoretical modeling, about 18% overall, driven by parton shower, hadronization, matrix-element/shower matching, PDF choice, and scale variations [2308.00716]. This experimentally established quark–gluon tagging as a calibrated analysis tool rather than a purely simulation-level concept.

## 4. Perturbative structure, Casimir scaling, and optimal observables

The analytically controlled starting point is the eikonal or double-logarithmic limit, where quark–gluon discrimination is governed solely by the initiating-parton color factor. For IRC-safe angularities $e_\beta \equiv \lambda^1_\beta$, the LL cumulative distributions obey Casimir scaling,
$$
\Sigma_i^{\rm LL}(e_\beta) = \exp\!\left[-\frac{\alpha_s C_i}{\pi \beta}\,\log^2 e_\beta\right], \qquad i=q,g,
$$
or equivalently
$$
\Sigma_q(\lambda)=e^{-C_F r(\lambda)},\qquad \Sigma_g(\lambda)=e^{-C_A r(\lambda)} .
$$
For Casimir-scaling observables the classifier separation becomes a universal number, with the QCD benchmark $\Delta_{\rm QCD}\simeq 0.1286$ [1704.03878]. This establishes the canonical perturbative baseline.

Beyond LL, the literature emphasizes two points. First, modern angularity calculations exhibit the first departures from pure Casimir scaling. In $Z+$jet production, jet angularities were computed at NLO+NLL$'$ using the Banfi–Salam–Zanderighi flavor-$k_t$ algorithm for IRC-safe jet-flavor assignment, and the resulting quark tag on the leading jet was used to enhance the initial-state gluon purity of the sample [2110.09337]. The leading-order purity
$$
f_g=\frac{\sigma_{qg}}{\sigma_{qq}+\sigma_{qg}}
$$
is promoted after tagging to
$$
\tilde f_g = \frac{\varepsilon_q\,\sigma_{qg}}{\varepsilon_g\,\sigma_{qq}+\varepsilon_q\,\sigma_{qg}} ,
$$
with $\varepsilon_q$ the quark efficiency and $\varepsilon_g$ the gluon mistag rate [2110.09337]. In PYTHIA, this tagging procedure improved the gluon purity by roughly 10%, with similar qualitative improvement after grooming.

Second, the likelihood-ratio viewpoint provides a systematic notion of optimality. In an independent-emission eikonal picture, the log-likelihood ratio for a jet with $M$ emissions reduces to
$$
\ln L_{q/g}^{\text{LL}} = M \ln\frac{C_F}{C_A},
$$
so multiplicity is optimal at this level [2207.12411]. Beyond the eikonal limit, the optimal observable becomes a linear combination of weighted multiplicities
$$
n^{(\kappa)}=\sum_{n=1}^M z_n^\kappa ,
$$
with coefficients determined by the splitting functions [2207.12411]. This result gives analytic support to the long-standing empirical importance of multiplicity-like observables.

The same analytic program also exposes the main limitation: quark–gluon separation is highly sensitive to higher-order perturbative effects and to hadronization [1704.03878]. Comparisons among Pythia, Herwig, Sherpa, Vincia, Deductor, Ariadne, and Dire showed substantial spread in predicted discrimination, and hadronization shape functions can materially change separation power even for IRC-safe observables [1704.03878]. The practical consequence is that calculability and robustness do not automatically coincide; they must be established jointly.

## 5. Machine learning, unsupervised inference, and interpretability

Machine-learning approaches were adopted early because low-level constituent information can capture correlations that are compressed away by hand-engineered summary variables. Recursive neural networks operating on the sequential-clustering tree outperformed a BDT by a few percent in gluon rejection rate, and the tree structure itself already carried much of the useful information for quark–gluon discrimination [1711.02633]. The LoLa architecture, built from constituent four-vectors and trainable Lorentz-layer combinations, also showed immediate benefit in benchmark mono-jet and di-jet-resonance applications, though detector effects reduced the margin over simpler multivariate baselines [1812.09223].

A central theoretical issue for low-level networks is safety. Comparing PFNs and EFNs, one study found that PFNs outperform EFNs on hadron-level jets, but the gap essentially disappears when hadronization is turned off [2103.09103]. The extra PFN performance was traced to IRC-unsafe information associated mainly with soft, narrow-angle structure induced by hadronization, and the same work showed that interpretable high-level observables can reproduce PFN performance at the 99% level or higher [2103.09103]. This provides a concrete route for systematic validation: one can measure the surrogate observables directly and test whether the network is exploiting well-modeled features.

A different direction is fully unsupervised tagging. For SoftDrop multiplicity $n_{\mathrm{SD}}$, quark- and gluon-initiated jets are approximately Poisson-distributed at leading-logarithmic accuracy, so a mixed sample can be modeled as a two-component Poisson mixture [2112.11352]. Maximum-likelihood or Bayesian inference then yields both mixture fractions and class-conditional Poisson rates directly from data, and the corresponding posterior responsibilities define the tagger. Reported accuracy was roughly $0.65$–$0.73$ on Pythia and $0.62$–$0.70$ on Herwig, with the Bayesian posterior-averaged tagger reaching about 0.71 accuracy [2112.11352]. Low KL divergence and low Hellinger distance correlated well with high classification accuracy, enabling unsupervised hyperparameter selection, and detector-inspired angular smearing did not significantly degrade performance [2112.11352].

Recent work has focused on interpretability and robustness. A latent-space study of ParticleNet-Lite showed that the first 3–5 principal components already recover essentially all performance, with the dominant directions corresponding to multiplicity and particle-type diversity, radial jet shape, and fragmentation or energy dispersion [2507.21214]. The same analysis found that standard SHAP can produce distorted attributions when inputs are correlated, and symbolic regression can approximate the tagger output with compact nonlinear formulas [2507.21214]. In parallel, a resilience study argued that decorrelation can fail in quark–gluon tagging because the most distinctive feature is aligned with theory uncertainty; it proposed conditional training on interpolated Pythia and Herwig samples with a controlled Bayesian ParticleNet-Lite as a more resilient framework [2212.10493]. Together, these results suggest that modern quark–gluon tagging is no longer evaluated only by AUC or rejection, but also by safety, calibration, and stability under generator variation.

## 6. Applications, current frontiers, and limitations

The applications are diverse because quark–gluon composition carries process information. In electroweak boson plus jet production, tagging the leading jet as quark-initiated preferentially selects the $qg\to qZ$ channel and thereby enhances the initial-state gluon component; this motivates tagged $Z$-boson transverse-momentum spectra as observables for probing the gluon PDF [2108.10024]. In invisible Higgs searches from gluon fusion, tagging the leading ISR jet as gluon-like was proposed as a way to separate signal from electroweak vector-boson backgrounds, with reported 95% CL upper limits on $\sigma\times \mathrm{BR}(H\to\mathrm{inv})/\sigma_{\rm SM}$ improving from $60.2^{+30.0}_{-18.3}\%$ using only $E_T^{\rm miss}$ to $20.4^{+10.1}_{-5.99}\%$ using girth only, $8.3^{+4.46}_{-2.55}\%$ using a DNN on jet substructure only, and $5.2^{+2.83}_{-1.54}\%$ using a DNN on all features [2003.06822].

The scope is not limited to hadron-collider final states. In Deep Inelastic Scattering at next-to-eikonal accuracy, back-to-back quark–gluon dijets induced by $t$-channel quark exchange factorize onto the unpolarized quark TMD $f_1^q(x=0,\mathbf{k})$, making heavy-flavor-tagged quark–gluon dijets a potential new probe at the Electron Ion Collider [2303.12691]. This is not conventional quark–gluon tagging in the LHC sense, but it uses the identifiable quark–gluon final state as a controlled handle on a quark-exchange process.

The current experimental frontier is constituent-level transformers. For HL-LHC conditions with 140 pile-up interactions, ATLAS studies using the Particle Transformer (ParT) on anti-$k_t$, $R=0.4$ jets reconstructed from PFOs found about 10% better gluon rejection at low $p_T$ and up to 25% improvement at high $p_T$ in the central region, together with 20–30% better gluon rejection in the forward region relative to fully connected baselines [2509.14759]. The performance remained stable as pile-up increased from 60 to 200 interactions, and the gains were tied directly to constituent, track, and topo-tower information plus the extended forward tracking of the ITk [2509.14759]. In Run 2 and Run 3 data, the ATLAS DeParT transformer operated over $p_T>20$ GeV and $|\eta|<4.5$, and a jet-topics calibration reduced systematic uncertainty by up to 20% in some phase-space regions relative to the matrix method [2512.03949].

The limitations are equally well established. Raw performance depends strongly on parton shower and hadronization modeling, on jet $p_T$, on pileup, and on jet region in $\eta$ [1409.3072]. Generator differences can be physically consequential even for IRC-safe observables, and the shower–hadronization boundary is itself ambiguous [1704.03878]. Topology dependence remains at the percent level for same-flavor jets and is sensitive to labeling schemes, grooming, jet radius, and the collider environment [1810.05653]. A plausible implication is that the field’s long-term direction is not toward a single universal tagger, but toward a family of calibrated, process-aware, and increasingly data-driven discriminants whose theoretical control is explicit.

Source: https://www.emergentmind.com/topics/quark-gluon-tagging