Papers
Topics
Authors
Recent
Search
2000 character limit reached

Optimal Observable Technique (OOT)

Updated 8 July 2026
  • Optimal Observable Technique is a parameter-estimation framework that decomposes differential distributions into known basis functions to minimize statistical uncertainty.
  • It is applied in collider phenomenology and quantum metrology to extract anomalous couplings, Wilson coefficients, and achieve precision by saturating the quantum Cramér–Rao bound.
  • Recent extensions, such as the Optimal Observable Machine, integrate machine learning and detector awareness to optimize the full inference chain while controlling unfolding bias.

Searching arXiv for recent and foundational papers on the Optimal Observable Technique and closely related formulations. The Optimal Observable Technique (OOT) denotes, in the literature considered here, a class of parameter-estimation methods in which a differential distribution is written in a basis of known phase-space functions and the data are weighted so that the coefficients of interest are extracted with minimum statistical uncertainty. In collider phenomenology, OOT is used to determine anomalous couplings, Wilson coefficients, or model parameters from the full shape of a cross section rather than from a total rate alone; in quantum metrology, an “optimal observable” is one whose error-propagation precision saturates the quantum Cramér–Rao bound. A recent machine-learning development, the Optimal Observable Machine (OOM), is closely aligned in spirit with OOT but optimizes the expected precision of a full detector-level and unfolding-aware measurement chain rather than an analytic event-by-event statistic (Hioki et al., 2013, Calcuttawala et al., 2017, Zhong et al., 2013, Mohr et al., 13 Jan 2026).

1. Statistical definition and canonical formulas

A standard OOT starting point is a decomposition of the theoretical distribution into known basis functions,

Σ(ϕ)=icifi(ϕ),\Sigma(\phi)=\sum_i c_i f_i(\phi),

or, in a generalized notation used in several collider applications,

O(ϕ)=dσdϕ=gifi(ϕ).\mathcal O(\phi)=\frac{d\sigma}{d\phi}=g_i f_i(\phi).

Here fi(ϕ)f_i(\phi) are known functions of the phase-space variable ϕ\phi, while cic_i or gig_i are coefficients to be extracted. An important point emphasized in comparative studies is that these coefficients need not be the Lagrangian couplings themselves; they may be nonlinear functions of the underlying new-physics parameters (Bhattacharya et al., 2023).

The coefficients are obtained from weighted integrals,

wi(ϕ)O(ϕ)dϕ=gi,\int w_i(\phi)\,\mathcal O(\phi)\,d\phi = g_i,

and the statistically optimal weights are

wi(ϕ)=jXijfj(ϕ)O(ϕ),Xij=(M1)ij,Mij=fi(ϕ)fj(ϕ)O(ϕ)dϕ.w_i(\phi)=\frac{\sum_j X_{ij}f_j(\phi)}{\mathcal O(\phi)}, \qquad X_{ij}=(M^{-1})_{ij}, \qquad M_{ij}=\int \frac{f_i(\phi)f_j(\phi)}{\mathcal O(\phi)}\,d\phi.

For this choice, the covariance matrix becomes

Vij=(M1)ijσTN,σT=O(ϕ)dϕ.V_{ij}=\frac{(M^{-1})_{ij}\sigma_T}{N}, \qquad \sigma_T=\int \mathcal O(\phi)\,d\phi.

When event counts are written as N=σ×L×ϵN=\sigma\times{\cal L}\times\epsilon, luminosity and efficiency enter directly through the same covariance scaling. This structure is the core of the method across flavor, top, lepton-collider, and heavy-fermion applications (Calcuttawala et al., 2017, Bhattacharya et al., 2021).

The same formalism is often coupled to a model-separation statistic,

O(ϕ)=dσdϕ=gifi(ϕ).\mathcal O(\phi)=\frac{d\sigma}{d\phi}=g_i f_i(\phi).0

or, in equivalent notation,

O(ϕ)=dσdϕ=gifi(ϕ).\mathcal O(\phi)=\frac{d\sigma}{d\phi}=g_i f_i(\phi).1

where O(ϕ)=dσdϕ=gifi(ϕ).\mathcal O(\phi)=\frac{d\sigma}{d\phi}=g_i f_i(\phi).2 or O(ϕ)=dσdϕ=gifi(ϕ).\mathcal O(\phi)=\frac{d\sigma}{d\phi}=g_i f_i(\phi).3 are seed values corresponding to a reference model. In the cited phenomenological studies, O(ϕ)=dσdϕ=gifi(ϕ).\mathcal O(\phi)=\frac{d\sigma}{d\phi}=g_i f_i(\phi).4 is interpreted as an O(ϕ)=dσdϕ=gifi(ϕ).\mathcal O(\phi)=\frac{d\sigma}{d\phi}=g_i f_i(\phi).5 separation from the reference point (Calcuttawala et al., 2017, Jahedi et al., 2024).

2. Classical high-energy-physics interpretation

In the classical or high-energy-physics sense, OOT seeks an observable whose distribution saturates local information about a parameter, often near a reference point, and this is tied to the score or to interference structures in the differential cross section. The score-based formulation is the familiar

O(ϕ)=dσdϕ=gifi(ϕ).\mathcal O(\phi)=\frac{d\sigma}{d\phi}=g_i f_i(\phi).6

while a local expansion may be written as

O(ϕ)=dσdϕ=gifi(ϕ).\mathcal O(\phi)=\frac{d\sigma}{d\phi}=g_i f_i(\phi).7

with the usual OOT ratio O(ϕ)=dσdϕ=gifi(ϕ).\mathcal O(\phi)=\frac{d\sigma}{d\phi}=g_i f_i(\phi).8 appearing as the locally optimal observable. This establishes the link between OOT and Fisher-information saturation at the level of the differential distribution (Mohr et al., 13 Jan 2026).

In practical collider analyses, however, OOT is often implemented not as an analytic score observable but as an optimal weighting of basis shapes. This distinction is important because many applications are not purely linear in the underlying couplings. In O(ϕ)=dσdϕ=gifi(ϕ).\mathcal O(\phi)=\frac{d\sigma}{d\phi}=g_i f_i(\phi).9 with selected semileptonic operators, for example, the cited comparative study emphasizes a regime in which there is no SM–EFT interference, so the OOT coefficients are quadratic combinations such as fi(ϕ)f_i(\phi)0, fi(ϕ)f_i(\phi)1, and fi(ϕ)f_i(\phi)2, even though the basis decomposition remains linear in those derived coefficients (Bhattacharya et al., 2023).

The method is therefore best understood as a minimum-variance inference framework for differential shapes. This suggests that “optimal observable” is not restricted to one formal expression: in some settings it is a score-like event variable, in others it is an optimal weight acting on a measured distribution, and in newer detector-aware constructions it becomes an optimized publishable observable rather than a closed-form sufficient statistic (Mohr et al., 13 Jan 2026).

3. Representative phenomenological applications

OOT has been applied to a wide range of collider and flavor problems in which a small set of physically interpretable differential structures carries most of the parameter sensitivity.

Domain Representative process OOT role
Top EFT at hadron colliders fi(ϕ)f_i(\phi)3 (Hioki et al., 2013) Extracts fi(ϕ)f_i(\phi)4 and fi(ϕ)f_i(\phi)5 from lepton angular and energy distributions
Rare flavor decays fi(ϕ)f_i(\phi)6 and fi(ϕ)f_i(\phi)7 (Calcuttawala et al., 2017) Optimizes fi(ϕ)f_i(\phi)8-shape information and polarization observables
Diboson production at lepton colliders fi(ϕ)f_i(\phi)9 (Jahedi et al., 2024) Constrains cTGC SMEFT coefficients using ϕ\phi0, ϕ\phi1-helicity channels, and beam polarization
Lepton-flavor violation ϕ\phi2 and ϕ\phi3 (Jahedi et al., 2024) Uses angular basis functions ϕ\phi4, ϕ\phi5, ϕ\phi6 to bound four-Fermi SMEFT couplings
Heavy charged fermions ϕ\phi7 (Bhattacharya et al., 2021) Measures vector and axial ϕ\phi8 couplings from production-angle shapes
Comparative regime study ϕ\phi9 and cic_i0 (Bhattacharya et al., 2023) Compares SM-dominated and NP-dominated OOT performance

Across these applications, several recurrent patterns emerge. First, OOT is most useful when the available observables are few but their shapes are still informative; this is explicit in the cic_i1 invisible analysis, where the cic_i2 dependence and polarization observables recover substantial discriminating power despite missing-energy final states (Calcuttawala et al., 2017). Second, beam polarization and final-state helicity resolution can materially improve the information matrix by changing both event yields and the relative size of the parameter-dependent terms, as shown for charged triple gauge couplings and LFV contact interactions (Jahedi et al., 2024, Jahedi et al., 2024). Third, comparative studies find that OOT can outperform a standard binned cic_i3 analysis most clearly when new physics is a small distortion on top of a large Standard Model background, whereas the gap narrows when the new-physics contribution dominates the event sample (Bhattacharya et al., 2023).

4. Quantum-metrological formulation

A technically distinct but conceptually related “optimal observable” problem appears in single-parameter quantum estimation. The relevant error-propagation formula is

cic_i4

while the ultimate quantum limit is the quantum Cramér–Rao bound

cic_i5

with symmetric logarithmic derivative (SLD) defined by

cic_i6

The optimality condition derived for error-propagation saturation is

cic_i7

which is necessary and sufficient for attaining the QCRB through measurement of cic_i8 (Zhong et al., 2013).

This result is closely analogous to the classical OOT relation between optimal observables and the score. In the quantum setting, the SLD is the analogue of the classical score, and the QFI plays the role of Fisher information. The centered observable cic_i9 must therefore coincide with the SLD on the support of the state, up to scaling and an additive constant. That is not the same formal object as the high-energy-physics event observable gig_i0, but it is the same local-information principle expressed at the operator level (Zhong et al., 2013).

For GHZ states in Ramsey interferometry, the analysis finds that the general optimal separable observable is

gig_i1

with gig_i2 and gig_i3 arbitrary real numbers not both zero. This family attains

gig_i4

and after a gig_i5 Ramsey rotation contains parity measurement as the case gig_i6. The same work also argues that the claim that separable measurements cannot achieve Heisenberg scaling for entangled states is incorrect in general: for GHZ states, separable measurements can attain that scaling except at isolated singular parameter points where the derivative of the measured expectation value vanishes (Zhong et al., 2013).

5. Detector-aware and machine-learned extension: the Optimal Observable Machine

The Optimal Observable Machine (OOM) is a machine-learning framework for constructing generator-level observables optimized for parameter extraction in particle-physics analyses. It is explicitly not presented as a direct reformulation of classical OOT, and it is not explicitly derived from the score gig_i7, from the expansion gig_i8, or from the usual ratio gig_i9. What it does do is learn an observable specifically to maximize parameter sensitivity in a full measurement pipeline that includes detector response, nuisance parameters, and unfolding-bias control (Mohr et al., 13 Jan 2026).

The formal setup introduces a learned generator-level observable

wi(ϕ)O(ϕ)dϕ=gi,\int w_i(\phi)\,\mathcal O(\phi)\,d\phi = g_i,0

and a detector-level distribution

wi(ϕ)O(ϕ)dϕ=gi,\int w_i(\phi)\,\mathcal O(\phi)\,d\phi = g_i,1

related, in the idealized no-background case, by

wi(ϕ)O(ϕ)dϕ=gi,\int w_i(\phi)\,\mathcal O(\phi)\,d\phi = g_i,2

The measurement objective is encoded in a profile likelihood,

wi(ϕ)O(ϕ)dϕ=gi,\int w_i(\phi)\,\mathcal O(\phi)\,d\phi = g_i,3

with Hessian

wi(ϕ)O(ϕ)dϕ=gi,\int w_i(\phi)\,\mathcal O(\phi)\,d\phi = g_i,4

and training loss

wi(ϕ)O(ϕ)dϕ=gi,\int w_i(\phi)\,\mathcal O(\phi)\,d\phi = g_i,5

The paper explicitly states that this “utilizes the Fisher information formalism for the case of the Asimov data set,” so the optimized quantity is expected profile-likelihood precision rather than the score itself (Mohr et al., 13 Jan 2026).

To propagate gradients through both generator- and detector-level networks, OOM defines

wi(ϕ)O(ϕ)dϕ=gi,\int w_i(\phi)\,\mathcal O(\phi)\,d\phi = g_i,6

and uses

wi(ϕ)O(ϕ)dϕ=gi,\int w_i(\phi)\,\mathcal O(\phi)\,d\phi = g_i,7

A further response-matrix penalty,

wi(ϕ)O(ϕ)dϕ=gi,\int w_i(\phi)\,\mathcal O(\phi)\,d\phi = g_i,8

suppresses parameter-dependent unfolding bias by favoring observables whose response matrix is approximately independent of the parameter of interest. In the demonstrated top-quark application, simultaneous generator- and detector-level training without this constraint yields high sensitivity but induces bias, particularly for the rarer pseudoscalar component; scanning 50 trainings for each wi(ϕ)O(ϕ)dϕ=gi,\int w_i(\phi)\,\mathcal O(\phi)\,d\phi = g_i,9, the study identifies wi(ϕ)=jXijfj(ϕ)O(ϕ),Xij=(M1)ij,Mij=fi(ϕ)fj(ϕ)O(ϕ)dϕ.w_i(\phi)=\frac{\sum_j X_{ij}f_j(\phi)}{\mathcal O(\phi)}, \qquad X_{ij}=(M^{-1})_{ij}, \qquad M_{ij}=\int \frac{f_i(\phi)f_j(\phi)}{\mathcal O(\phi)}\,d\phi.0 as a good compromise between expected uncertainty and low bias (Mohr et al., 13 Jan 2026).

Relative to classical OOT, the shift is therefore from local analytic optimality at the differential-cross-section level to global optimization of the complete inference chain. The learned object is still low-dimensional and parameter-targeted, but it is also detector-aware, unfolding-aware, and explicitly motivated by long-term reinterpretability of unfolded generator-level measurements (Mohr et al., 13 Jan 2026).

6. Scope, limitations, and recurrent misconceptions

A persistent misconception is that OOT is synonymous with a single formula or a single research domain. The cited literature instead supports a broader view: classical collider OOT, score-like local-optimal constructions, SLD-based quantum-optimal observables, and OOM-style detector-aware learned observables all instantiate the same general idea of preserving maximal parameter sensitivity in a restricted observable space, but they are not formally identical (Zhong et al., 2013, Mohr et al., 13 Jan 2026).

Another common simplification is to treat OOT as automatically global and model-independent. In practice, many implementations are local or seed-dependent because the covariance matrix depends on the denominator wi(ϕ)=jXijfj(ϕ)O(ϕ),Xij=(M1)ij,Mij=fi(ϕ)fj(ϕ)O(ϕ)dϕ.w_i(\phi)=\frac{\sum_j X_{ij}f_j(\phi)}{\mathcal O(\phi)}, \qquad X_{ij}=(M^{-1})_{ij}, \qquad M_{ij}=\int \frac{f_i(\phi)f_j(\phi)}{\mathcal O(\phi)}\,d\phi.1, hence on the assumed parameter point. The wi(ϕ)=jXijfj(ϕ)O(ϕ),Xij=(M1)ij,Mij=fi(ϕ)fj(ϕ)O(ϕ)dϕ.w_i(\phi)=\frac{\sum_j X_{ij}f_j(\phi)}{\mathcal O(\phi)}, \qquad X_{ij}=(M^{-1})_{ij}, \qquad M_{ij}=\int \frac{f_i(\phi)f_j(\phi)}{\mathcal O(\phi)}\,d\phi.2 invisible analysis explicitly notes seed dependence, and the comparative study emphasizes that even in an SM-dominated regime it may be inaccurate to replace the full wi(ϕ)=jXijfj(ϕ)O(ϕ),Xij=(M1)ij,Mij=fi(ϕ)fj(ϕ)O(ϕ)dϕ.w_i(\phi)=\frac{\sum_j X_{ij}f_j(\phi)}{\mathcal O(\phi)}, \qquad X_{ij}=(M^{-1})_{ij}, \qquad M_{ij}=\int \frac{f_i(\phi)f_j(\phi)}{\mathcal O(\phi)}\,d\phi.3 by the SM distribution in the information matrix if the new-physics contribution is not negligible (Calcuttawala et al., 2017, Bhattacharya et al., 2023).

A further limitation is that many phenomenological OOT studies remain statistical projections rather than full experimental forecasts. The top-coupling analysis at the LHC includes only statistical errors and reports numerical instability in the inversion of the OOT matrices, reflecting parameter correlations and near-degeneracies (Hioki et al., 2013). The cTGC study at wi(ϕ)=jXijfj(ϕ)O(ϕ),Xij=(M1)ij,Mij=fi(ϕ)fj(ϕ)O(ϕ)dϕ.w_i(\phi)=\frac{\sum_j X_{ij}f_j(\phi)}{\mathcal O(\phi)}, \qquad X_{ij}=(M^{-1})_{ij}, \qquad M_{ij}=\int \frac{f_i(\phi)f_j(\phi)}{\mathcal O(\phi)}\,d\phi.4 colliders is likewise essentially statistical and does not include a realistic detector simulation or systematic uncertainties in the OOT framework (Jahedi et al., 2024). The LFV study uses detector-level cut efficiencies and discusses the role of wi(ϕ)=jXijfj(ϕ)O(ϕ),Xij=(M1)ij,Mij=fi(ϕ)fj(ϕ)O(ϕ)dϕ.w_i(\phi)=\frac{\sum_j X_{ij}f_j(\phi)}{\mathcal O(\phi)}, \qquad X_{ij}=(M^{-1})_{ij}, \qquad M_{ij}=\int \frac{f_i(\phi)f_j(\phi)}{\mathcal O(\phi)}\,d\phi.5 and wi(ϕ)=jXijfj(ϕ)O(ϕ),Xij=(M1)ij,Mij=fi(ϕ)fj(ϕ)O(ϕ)dϕ.w_i(\phi)=\frac{\sum_j X_{ij}f_j(\phi)}{\mathcal O(\phi)}, \qquad X_{ij}=(M^{-1})_{ij}, \qquad M_{ij}=\int \frac{f_i(\phi)f_j(\phi)}{\mathcal O(\phi)}\,d\phi.6, but does not provide an explicit background-inclusive modification of the OOT covariance formulas (Jahedi et al., 2024). The heavy-charged-fermion analysis folds acceptance and background rejection into an overall efficiency factor wi(ϕ)=jXijfj(ϕ)O(ϕ),Xij=(M1)ij,Mij=fi(ϕ)fj(ϕ)O(ϕ)dϕ.w_i(\phi)=\frac{\sum_j X_{ij}f_j(\phi)}{\mathcal O(\phi)}, \qquad X_{ij}=(M^{-1})_{ij}, \qquad M_{ij}=\int \frac{f_i(\phi)f_j(\phi)}{\mathcal O(\phi)}\,d\phi.7 rather than a full detector-level likelihood (Bhattacharya et al., 2021).

Finally, OOT cannot compensate for a lack of parameter identifiability in the observable itself. The wi(ϕ)=jXijfj(ϕ)O(ϕ),Xij=(M1)ij,Mij=fi(ϕ)fj(ϕ)O(ϕ)dϕ.w_i(\phi)=\frac{\sum_j X_{ij}f_j(\phi)}{\mathcal O(\phi)}, \qquad X_{ij}=(M^{-1})_{ij}, \qquad M_{ij}=\int \frac{f_i(\phi)f_j(\phi)}{\mathcal O(\phi)}\,d\phi.8 channel under LFU/no-LFV assumptions is explicitly noted to be unsuitable for a genuine multi-parameter optimal-observable extraction because it depends only on the single combination wi(ϕ)=jXijfj(ϕ)O(ϕ),Xij=(M1)ij,Mij=fi(ϕ)fj(ϕ)O(ϕ)dϕ.w_i(\phi)=\frac{\sum_j X_{ij}f_j(\phi)}{\mathcal O(\phi)}, \qquad X_{ij}=(M^{-1})_{ij}, \qquad M_{ij}=\int \frac{f_i(\phi)f_j(\phi)}{\mathcal O(\phi)}\,d\phi.9 (Calcuttawala et al., 2017). By contrast, machine-learned extensions such as OOM can incorporate detector effects, nuisance profiling, and unfolding-bias control directly into the optimization, but they remain parameter-specific, simulation-dependent, and subject to a precision-versus-bias trade-off governed by the response-matrix regularization (Mohr et al., 13 Jan 2026).

In this sense, OOT is best regarded not as a fixed algorithm but as an inference principle: choose or construct the observable representation that minimizes the expected uncertainty on the parameter of interest, subject to the structural constraints of the measurement problem.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Optimal Observable Technique (OOT).