TH-Score: A Multifaceted Performance Metric
- TH-Score is a versatile performance metric with context-specific definitions across fields such as linear-system identification, survival prediction, tensor analysis, and nuclear spectroscopy.
- It quantifies performance through threshold-based evaluations, threshold-free summaries, and structural graph measures, enabling effective order recovery and error control.
- The metric offers actionable insights by compressing complex trade-offs into scalar values, guiding comparative analyses in speech synthesis, event detection, and nuclear-optical applications.
Searching arXiv for the cited papers to ground the article. TH-Score is not a single standardized quantity in the arXiv literature surveyed here. Instead, the label appears as an inferred or contextual name for several distinct scalar performance measures: a rank-adaptive score for Thresholded Ho–Kalman in linear-system identification, threshold-based or threshold-free summaries in event-detection and censored survival evaluation, a structural score for tensor-train core orderings in stochastic differential equations, and, by naming overlap rather than identity, TTScore for speech-synthesis evaluation. In thorium-229 research, the same label can only be read metaphorically as a figure of merit for excitation, spectroscopy, or clock performance rather than as an author-defined variable (Zheng et al., 10 Jun 2025, Kinjo et al., 29 Mar 2025, Yuan et al., 2016, Ulgen et al., 24 Sep 2025, Porsev et al., 2010, Beloy, 2014, Shigekawa et al., 2021).
1. Terminological status and domain-specific meanings
The available works do not use “TH-Score” in a single uniform sense. In some cases, the term is not explicit in the paper but can be grounded in the paper’s main quantitative guarantees; in others, a closely related name denotes a formally defined metric. This suggests that TH-Score is best treated as a domain-dependent label whose meaning must be fixed locally by the underlying model, loss, or physical observable.
| Context | Formal object | Role |
|---|---|---|
| Rank-adaptive LTI identification | Frobenius Hankel error, order recovery, Markov-parameter error | Thresholded Ho–Kalman performance |
| Sound event detection | Metric at an operating threshold selected from all-threshold curves | Threshold-based evaluation derived from threshold-independent computation |
| Censored survival prediction | Threshold-free average PPV | |
| TT ordering for SDE moments | Weighted cut score across the TT mid-bond | |
| Speech synthesis | TTScore-int, TTScore-pro | Text-conditioned token-likelihood evaluation |
| Th nuclear physics | , , , | Performance figures for excitation, spectroscopy, and clocking |
A common misconception is that TH-Score names a universally accepted benchmark. The literature summarized here does not support that interpretation. Several papers explicitly state that the term is not used by the authors, and the score-like object must instead be reconstructed from the paper’s proven bounds or experimentally central observables (Zheng et al., 10 Jun 2025, Porsev et al., 2010, Beloy, 2014, Shigekawa et al., 2021).
2. Rank-adaptive TH-Score in Thresholded Ho–Kalman
In linear time-invariant identification, the natural TH-Score arises from the paper “Minimal Order Recovery through Rank-adaptive Identification” (Zheng et al., 10 Jun 2025). The setting is the discrete-time partially observed LTI system
with unknown , i.i.d. Gaussian inputs , measurement noise 0, unknown system order 1, and Markov parameters 2. The finite Hankel matrix is
3
and under 4 it has rank exactly 5 (Zheng et al., 10 Jun 2025).
Thresholded Ho–Kalman (THK) first estimates a Hankel-like matrix 6 by least squares and then applies singular-value thresholding,
7
The effective rank 8 is used as the estimated system order, after which Ho–Kalman is run on 9 to recover 0 up to similarity (Zheng et al., 10 Jun 2025).
The central rank-adaptive guarantee is the Frobenius bound
1
valid whenever 2 for 3. This inequality makes explicit the variance-like term 4 and the bias-like tail 5. The same analysis shows 6 when 7, and exact order recovery 8 when additionally 9 (Zheng et al., 10 Jun 2025).
A principled TH-Score in this setting is therefore naturally composite. Its first component is Hankel recovery,
0
its second is the probability of correct order recovery,
1
and its third is the induced Markov-parameter error, controlled by the Hankel error through finite-time Ho–Kalman perturbation bounds. The paper further proposes a clean theoretical scalarization,
2
which combines rank-adaptive efficiency with an order-recovery penalty. This is not an author-defined quantity, but it is fully grounded in the paper’s finite-sample analysis (Zheng et al., 10 Jun 2025).
An important significance claim is comparative rather than absolute: the sample complexity and finite-sample error rates match those of state-of-the-art methods that assume prior knowledge of the system order. In that sense, the THK interpretation of TH-Score measures the cost of order ignorance and shows that, under the proved thresholds, this cost can be asymptotically eliminated (Zheng et al., 10 Jun 2025).
3. Threshold dependence and threshold-free summaries
In sound event detection, “TH-Score” is best understood as a threshold-based score extracted from a threshold-independent evaluation pipeline rather than as a fixed metric name. The paper “Threshold Independent Evaluation of Sound Event Detection Scores” (Ebbers et al., 2022) treats continuous framewise scores 3, converts them to events by thresholding, and shows that conventional event-based metrics depend strongly on the choice of threshold 4. Its core technical contribution is exact computation of collar-based and intersection-based operating curves for all effectively distinct thresholds by accumulating per-score deltas such as 5, 6, and 7. This yields threshold-independent PSD-ROC and PSDS computation without sparse threshold sweeps.
Within that framework, a TH-Score can be defined as any threshold-based operating-point metric selected from the full curve. For example, if the chosen score is collar-based 8,
9
then 0. The paper reports that, on the evaluation set, moving from fixed 1 to thresholds tuned on the validation PR curve improves F1 from 2 to 3. It also reports exact PSDS 4, versus approximately 5 for 50 linear thresholds and approximately 6 for 500 linear thresholds, showing that coarse threshold grids can significantly underestimate threshold-independent scores (Ebbers et al., 2022).
A related but conceptually distinct threshold-free construction appears in censored survival prediction. The paper “A Threshold-free Prospective Prediction Accuracy Measure for Censored Time to Event Data” (Yuan et al., 2016) defines the time-dependent PPV and TPF
7
and then introduces the threshold-free summary
8
the time-dependent average positive predictive value. Here 9 is the area under the time-dependent precision–recall curve and lies in 0, with 1. The lower endpoint corresponds to a non-informative score and the upper endpoint to perfect separation (Yuan et al., 2016).
The contrast is instructive. In the SED setting, TH-Score is naturally a threshold-selected score derived from an exact all-threshold evaluator. In the censored-survival setting, the analogous object is explicitly threshold-free and averages PPV over all thresholds induced by the case-score distribution. The shared conceptual core is not a common formula but a common concern: replacing arbitrary threshold choice by either exact operating-curve optimization or principled integration over thresholds (Ebbers et al., 2022, Yuan et al., 2016).
4. Structural TH-Score for tensor-train orderings in SDE moment computation
A formally defined score appears in the paper “Permutation of Tensor-Train Cores for Computing Moments on Stochastic Differential Equations” (Kinjo et al., 29 Mar 2025). The problem is TT-based solution of dual ODE systems arising from stochastic differential equations, where the ordering of TT cores strongly affects numerical accuracy. For a dimension 2, interaction matrix 3, and permutation 4 specifying the TT-core order, the score is
5
where 6 when species 7 and 8 lie on opposite sides of the central TT cut and 9 otherwise. For even 0, the cut is 1; for odd 2, the paper computes the two central cuts 3 and 4 and averages the resulting scores (Kinjo et al., 29 Mar 2025).
This score is a weighted cut value on the interaction graph induced by 5. Its conceptual meaning is that strongly interacting species should not be separated by the middle TT bond if one aims to keep TT ranks small and truncation error low. A low score means that most strong couplings remain within the left or right half of the TT chain; a high score means that many strong interactions are cut by the central bond. The score is purely structural: it depends only on the ordering and the SDE interaction matrix, not on the moments being computed or on the observed numerical error (Kinjo et al., 29 Mar 2025).
The paper’s numerical evidence shows a positive correlation between score and relative error. For cascade models with nearest-neighbor interactions and for random interaction matrices, mean relative error increases with score, although the scatter grows with dimension and with more complex random couplings. The authors do not prove an error bound in terms of 6, but they conclude that orderings minimizing the score tend to yield higher accuracy (Kinjo et al., 29 Mar 2025).
In this context, TH-Score is not threshold-related at all. It is instead a graph-theoretic surrogate for TT representability. That distinction is important because it prevents conflation with threshold-free or threshold-selected scores in evaluation problems. The commonality is only that each score compresses a high-dimensional performance or complexity trade-off into a single scalar functional (Kinjo et al., 29 Mar 2025).
5. Thorium-229 figures of merit
In three thorium-229 papers, “TH-Score” is not an author-defined symbol but can be read as a figure of merit for excitation efficiency, spectroscopic sensitivity, or clock performance. The paper “Excitation of the isomeric 7Th nuclear state via an electronic bridge process in 8Th9” (Porsev et al., 2010) analyzes a two-photon electronic-bridge scheme in singly ionized thorium. Its central scalar performance quantity is the enhancement factor
0
which compares the EB-mediated nuclear transition rate to the bare nuclear M1 rate. For the favored 1 case, the paper finds 2, 3, and, with specified conventional laser parameters, an induced EB excitation rate
4
In that setting, a thorium TH-Score is naturally identified with EB enhancement and induced excitation rate, because those observables quantify how practical the scheme is for locating and driving the isomeric transition (Porsev et al., 2010).
A second thorium figure of merit appears in “Hyperfine structure in 5Th6 as a probe of the 7Th 8 9Th nuclear excitation energy” (Beloy, 2014). There the relevant quantities are the hyperfine-mixing parameters
0
which encode the strength of coupling between the nuclear ground state and isomer through the hyperfine interaction. Precision microwave spectroscopy of the 1 hyperfine manifolds, combined with atomic-structure calculations, is proposed as a route to extracting 2 from 3 and 4 from 5. Together with the isomer M1 decay rate, these parameters allow an indirect estimate of the nuclear excitation energy 6. The paper argues that realistic uncertainties could allow determination of 7 to about 8, so here the thorium TH-Score is best interpreted as spectroscopic leverage on the isomer energy (Beloy, 2014).
A third thorium figure of merit is explicit clock performance. The paper “Estimation of radiative half-life of 9Th by half-life measurement of other nuclear excited states in 0Th” (Shigekawa et al., 2021) measures half-lives of several low-lying states, uses the Alaga rule to infer
1
and obtains the radiative half-life
2
The same paper states that this corresponds to a relative natural linewidth
3
which it describes as indicating superb performance for an optical nuclear clock. In that usage, the thorium TH-Score is naturally the combination of 4, natural linewidth, and the feasibility of suppressing internal conversion so that the radiative transition becomes operationally accessible (Shigekawa et al., 2021).
These three thorium usages are related but non-identical. One emphasizes excitation rate, one spectroscopic inference of 5, and one clock-linewidth performance. The common feature is that each compresses the value of 6Th as a controllable nuclear-optical system into a small number of experimentally or theoretically central scalars (Porsev et al., 2010, Beloy, 2014, Shigekawa et al., 2021).
6. TTScore in speech synthesis and the problem of near-homonymous metrics
A near-homonymous but distinct metric family appears in “Objective Evaluation of Prosody and Intelligibility in Speech Synthesis via Conditional Prediction of Discrete Tokens” (Ulgen et al., 24 Sep 2025). The paper does not define TH-Score. It defines TTScore, with two components: 7 for intelligibility via content tokens, and
8
for prosody via phoneme-level prosody tokens. Both are reference-free, text-conditioned, token-likelihood-based scores (Ulgen et al., 24 Sep 2025).
The distinction from TH-Score matters because the naming overlap can obscure the metric’s actual semantics. TTScore-int is a conditional content-likelihood score and TTScore-pro a conditional prosody-likelihood score; neither is threshold-based, threshold-free in the survival-analysis sense, nor related to TT-core ordering. The paper reports that TTScore-int and TTScore-pro achieve stronger correlations with human judgments than conventional intelligibility or prosody-focused baselines on SOMOS, VoiceMOS, and TTSArena, but these are properties of the TTScore family specifically (Ulgen et al., 24 Sep 2025).
This near-homonymy clarifies a broader point about TH-Score as an encyclopedia topic. The label is not a unified technical object but a recurrent score metaphor instantiated differently across fields. In system identification it names a rank-adaptive oracle-style trade-off; in evaluation theory it concerns threshold choice or threshold elimination; in tensor methods it is a weighted cut on an interaction graph; in thorium-229 physics it denotes performance figures for excitation, spectroscopy, and clocking; and in speech synthesis the relevant formalism is not TH-Score at all but TTScore (Zheng et al., 10 Jun 2025, Yuan et al., 2016, Kinjo et al., 29 Mar 2025, Shigekawa et al., 2021, Ulgen et al., 24 Sep 2025).