Papers
Topics
Authors
Recent
Search
2000 character limit reached

TH-Score: A Multifaceted Performance Metric

Updated 17 July 2026
  • TH-Score is a versatile performance metric with context-specific definitions across fields such as linear-system identification, survival prediction, tensor analysis, and nuclear spectroscopy.
  • It quantifies performance through threshold-based evaluations, threshold-free summaries, and structural graph measures, enabling effective order recovery and error control.
  • The metric offers actionable insights by compressing complex trade-offs into scalar values, guiding comparative analyses in speech synthesis, event detection, and nuclear-optical applications.

Searching arXiv for the cited papers to ground the article. TH-Score is not a single standardized quantity in the arXiv literature surveyed here. Instead, the label appears as an inferred or contextual name for several distinct scalar performance measures: a rank-adaptive score for Thresholded Ho–Kalman in linear-system identification, threshold-based or threshold-free summaries in event-detection and censored survival evaluation, a structural score for tensor-train core orderings in stochastic differential equations, and, by naming overlap rather than identity, TTScore for speech-synthesis evaluation. In thorium-229 research, the same label can only be read metaphorically as a figure of merit for excitation, spectroscopy, or clock performance rather than as an author-defined variable (Zheng et al., 10 Jun 2025, Kinjo et al., 29 Mar 2025, Yuan et al., 2016, Ulgen et al., 24 Sep 2025, Porsev et al., 2010, Beloy, 2014, Shigekawa et al., 2021).

1. Terminological status and domain-specific meanings

The available works do not use “TH-Score” in a single uniform sense. In some cases, the term is not explicit in the paper but can be grounded in the paper’s main quantitative guarantees; in others, a closely related name denotes a formally defined metric. This suggests that TH-Score is best treated as a domain-dependent label whose meaning must be fixed locally by the underlying model, loss, or physical observable.

Context Formal object Role
Rank-adaptive LTI identification Frobenius Hankel error, order recovery, Markov-parameter error Thresholded Ho–Kalman performance
Sound event detection Metric at an operating threshold selected from all-threshold curves Threshold-based evaluation derived from threshold-independent computation
Censored survival prediction APt0AP_{t_0} Threshold-free average PPV
TT ordering for SDE moments sσs_\sigma Weighted cut score across the TT mid-bond
Speech synthesis TTScore-int, TTScore-pro Text-conditioned token-likelihood evaluation
229^{229}Th nuclear physics βM1\beta_{M1}, WEBinW^{\rm in}_{\rm EB}, ηk\eta_k, T1/2radT_{1/2}^{\rm rad} Performance figures for excitation, spectroscopy, and clocking

A common misconception is that TH-Score names a universally accepted benchmark. The literature summarized here does not support that interpretation. Several papers explicitly state that the term is not used by the authors, and the score-like object must instead be reconstructed from the paper’s proven bounds or experimentally central observables (Zheng et al., 10 Jun 2025, Porsev et al., 2010, Beloy, 2014, Shigekawa et al., 2021).

2. Rank-adaptive TH-Score in Thresholded Ho–Kalman

In linear time-invariant identification, the natural TH-Score arises from the paper “Minimal Order Recovery through Rank-adaptive Identification” (Zheng et al., 10 Jun 2025). The setting is the discrete-time partially observed LTI system

xt+1=Axt+But, yt=Cxt+zt,\begin{aligned} x_{t+1} &= A x_t + B u_t,\ y_t &= C x_t + z_t, \end{aligned}

with unknown A,B,CA,B,C, i.i.d. Gaussian inputs utN(0,σu2I)u_t\sim \mathcal N(0,\sigma_u^2 I), measurement noise sσs_\sigma0, unknown system order sσs_\sigma1, and Markov parameters sσs_\sigma2. The finite Hankel matrix is

sσs_\sigma3

and under sσs_\sigma4 it has rank exactly sσs_\sigma5 (Zheng et al., 10 Jun 2025).

Thresholded Ho–Kalman (THK) first estimates a Hankel-like matrix sσs_\sigma6 by least squares and then applies singular-value thresholding,

sσs_\sigma7

The effective rank sσs_\sigma8 is used as the estimated system order, after which Ho–Kalman is run on sσs_\sigma9 to recover 229^{229}0 up to similarity (Zheng et al., 10 Jun 2025).

The central rank-adaptive guarantee is the Frobenius bound

229^{229}1

valid whenever 229^{229}2 for 229^{229}3. This inequality makes explicit the variance-like term 229^{229}4 and the bias-like tail 229^{229}5. The same analysis shows 229^{229}6 when 229^{229}7, and exact order recovery 229^{229}8 when additionally 229^{229}9 (Zheng et al., 10 Jun 2025).

A principled TH-Score in this setting is therefore naturally composite. Its first component is Hankel recovery,

βM1\beta_{M1}0

its second is the probability of correct order recovery,

βM1\beta_{M1}1

and its third is the induced Markov-parameter error, controlled by the Hankel error through finite-time Ho–Kalman perturbation bounds. The paper further proposes a clean theoretical scalarization,

βM1\beta_{M1}2

which combines rank-adaptive efficiency with an order-recovery penalty. This is not an author-defined quantity, but it is fully grounded in the paper’s finite-sample analysis (Zheng et al., 10 Jun 2025).

An important significance claim is comparative rather than absolute: the sample complexity and finite-sample error rates match those of state-of-the-art methods that assume prior knowledge of the system order. In that sense, the THK interpretation of TH-Score measures the cost of order ignorance and shows that, under the proved thresholds, this cost can be asymptotically eliminated (Zheng et al., 10 Jun 2025).

3. Threshold dependence and threshold-free summaries

In sound event detection, “TH-Score” is best understood as a threshold-based score extracted from a threshold-independent evaluation pipeline rather than as a fixed metric name. The paper “Threshold Independent Evaluation of Sound Event Detection Scores” (Ebbers et al., 2022) treats continuous framewise scores βM1\beta_{M1}3, converts them to events by thresholding, and shows that conventional event-based metrics depend strongly on the choice of threshold βM1\beta_{M1}4. Its core technical contribution is exact computation of collar-based and intersection-based operating curves for all effectively distinct thresholds by accumulating per-score deltas such as βM1\beta_{M1}5, βM1\beta_{M1}6, and βM1\beta_{M1}7. This yields threshold-independent PSD-ROC and PSDS computation without sparse threshold sweeps.

Within that framework, a TH-Score can be defined as any threshold-based operating-point metric selected from the full curve. For example, if the chosen score is collar-based βM1\beta_{M1}8,

βM1\beta_{M1}9

then WEBinW^{\rm in}_{\rm EB}0. The paper reports that, on the evaluation set, moving from fixed WEBinW^{\rm in}_{\rm EB}1 to thresholds tuned on the validation PR curve improves F1 from WEBinW^{\rm in}_{\rm EB}2 to WEBinW^{\rm in}_{\rm EB}3. It also reports exact PSDS WEBinW^{\rm in}_{\rm EB}4, versus approximately WEBinW^{\rm in}_{\rm EB}5 for 50 linear thresholds and approximately WEBinW^{\rm in}_{\rm EB}6 for 500 linear thresholds, showing that coarse threshold grids can significantly underestimate threshold-independent scores (Ebbers et al., 2022).

A related but conceptually distinct threshold-free construction appears in censored survival prediction. The paper “A Threshold-free Prospective Prediction Accuracy Measure for Censored Time to Event Data” (Yuan et al., 2016) defines the time-dependent PPV and TPF

WEBinW^{\rm in}_{\rm EB}7

and then introduces the threshold-free summary

WEBinW^{\rm in}_{\rm EB}8

the time-dependent average positive predictive value. Here WEBinW^{\rm in}_{\rm EB}9 is the area under the time-dependent precision–recall curve and lies in ηk\eta_k0, with ηk\eta_k1. The lower endpoint corresponds to a non-informative score and the upper endpoint to perfect separation (Yuan et al., 2016).

The contrast is instructive. In the SED setting, TH-Score is naturally a threshold-selected score derived from an exact all-threshold evaluator. In the censored-survival setting, the analogous object is explicitly threshold-free and averages PPV over all thresholds induced by the case-score distribution. The shared conceptual core is not a common formula but a common concern: replacing arbitrary threshold choice by either exact operating-curve optimization or principled integration over thresholds (Ebbers et al., 2022, Yuan et al., 2016).

4. Structural TH-Score for tensor-train orderings in SDE moment computation

A formally defined score appears in the paper “Permutation of Tensor-Train Cores for Computing Moments on Stochastic Differential Equations” (Kinjo et al., 29 Mar 2025). The problem is TT-based solution of dual ODE systems arising from stochastic differential equations, where the ordering of TT cores strongly affects numerical accuracy. For a dimension ηk\eta_k2, interaction matrix ηk\eta_k3, and permutation ηk\eta_k4 specifying the TT-core order, the score is

ηk\eta_k5

where ηk\eta_k6 when species ηk\eta_k7 and ηk\eta_k8 lie on opposite sides of the central TT cut and ηk\eta_k9 otherwise. For even T1/2radT_{1/2}^{\rm rad}0, the cut is T1/2radT_{1/2}^{\rm rad}1; for odd T1/2radT_{1/2}^{\rm rad}2, the paper computes the two central cuts T1/2radT_{1/2}^{\rm rad}3 and T1/2radT_{1/2}^{\rm rad}4 and averages the resulting scores (Kinjo et al., 29 Mar 2025).

This score is a weighted cut value on the interaction graph induced by T1/2radT_{1/2}^{\rm rad}5. Its conceptual meaning is that strongly interacting species should not be separated by the middle TT bond if one aims to keep TT ranks small and truncation error low. A low score means that most strong couplings remain within the left or right half of the TT chain; a high score means that many strong interactions are cut by the central bond. The score is purely structural: it depends only on the ordering and the SDE interaction matrix, not on the moments being computed or on the observed numerical error (Kinjo et al., 29 Mar 2025).

The paper’s numerical evidence shows a positive correlation between score and relative error. For cascade models with nearest-neighbor interactions and for random interaction matrices, mean relative error increases with score, although the scatter grows with dimension and with more complex random couplings. The authors do not prove an error bound in terms of T1/2radT_{1/2}^{\rm rad}6, but they conclude that orderings minimizing the score tend to yield higher accuracy (Kinjo et al., 29 Mar 2025).

In this context, TH-Score is not threshold-related at all. It is instead a graph-theoretic surrogate for TT representability. That distinction is important because it prevents conflation with threshold-free or threshold-selected scores in evaluation problems. The commonality is only that each score compresses a high-dimensional performance or complexity trade-off into a single scalar functional (Kinjo et al., 29 Mar 2025).

5. Thorium-229 figures of merit

In three thorium-229 papers, “TH-Score” is not an author-defined symbol but can be read as a figure of merit for excitation efficiency, spectroscopic sensitivity, or clock performance. The paper “Excitation of the isomeric T1/2radT_{1/2}^{\rm rad}7Th nuclear state via an electronic bridge process in T1/2radT_{1/2}^{\rm rad}8ThT1/2radT_{1/2}^{\rm rad}9” (Porsev et al., 2010) analyzes a two-photon electronic-bridge scheme in singly ionized thorium. Its central scalar performance quantity is the enhancement factor

xt+1=Axt+But, yt=Cxt+zt,\begin{aligned} x_{t+1} &= A x_t + B u_t,\ y_t &= C x_t + z_t, \end{aligned}0

which compares the EB-mediated nuclear transition rate to the bare nuclear M1 rate. For the favored xt+1=Axt+But, yt=Cxt+zt,\begin{aligned} x_{t+1} &= A x_t + B u_t,\ y_t &= C x_t + z_t, \end{aligned}1 case, the paper finds xt+1=Axt+But, yt=Cxt+zt,\begin{aligned} x_{t+1} &= A x_t + B u_t,\ y_t &= C x_t + z_t, \end{aligned}2, xt+1=Axt+But, yt=Cxt+zt,\begin{aligned} x_{t+1} &= A x_t + B u_t,\ y_t &= C x_t + z_t, \end{aligned}3, and, with specified conventional laser parameters, an induced EB excitation rate

xt+1=Axt+But, yt=Cxt+zt,\begin{aligned} x_{t+1} &= A x_t + B u_t,\ y_t &= C x_t + z_t, \end{aligned}4

In that setting, a thorium TH-Score is naturally identified with EB enhancement and induced excitation rate, because those observables quantify how practical the scheme is for locating and driving the isomeric transition (Porsev et al., 2010).

A second thorium figure of merit appears in “Hyperfine structure in xt+1=Axt+But, yt=Cxt+zt,\begin{aligned} x_{t+1} &= A x_t + B u_t,\ y_t &= C x_t + z_t, \end{aligned}5Thxt+1=Axt+But, yt=Cxt+zt,\begin{aligned} x_{t+1} &= A x_t + B u_t,\ y_t &= C x_t + z_t, \end{aligned}6 as a probe of the xt+1=Axt+But, yt=Cxt+zt,\begin{aligned} x_{t+1} &= A x_t + B u_t,\ y_t &= C x_t + z_t, \end{aligned}7Th xt+1=Axt+But, yt=Cxt+zt,\begin{aligned} x_{t+1} &= A x_t + B u_t,\ y_t &= C x_t + z_t, \end{aligned}8 xt+1=Axt+But, yt=Cxt+zt,\begin{aligned} x_{t+1} &= A x_t + B u_t,\ y_t &= C x_t + z_t, \end{aligned}9Th nuclear excitation energy” (Beloy, 2014). There the relevant quantities are the hyperfine-mixing parameters

A,B,CA,B,C0

which encode the strength of coupling between the nuclear ground state and isomer through the hyperfine interaction. Precision microwave spectroscopy of the A,B,CA,B,C1 hyperfine manifolds, combined with atomic-structure calculations, is proposed as a route to extracting A,B,CA,B,C2 from A,B,CA,B,C3 and A,B,CA,B,C4 from A,B,CA,B,C5. Together with the isomer M1 decay rate, these parameters allow an indirect estimate of the nuclear excitation energy A,B,CA,B,C6. The paper argues that realistic uncertainties could allow determination of A,B,CA,B,C7 to about A,B,CA,B,C8, so here the thorium TH-Score is best interpreted as spectroscopic leverage on the isomer energy (Beloy, 2014).

A third thorium figure of merit is explicit clock performance. The paper “Estimation of radiative half-life of A,B,CA,B,C9Th by half-life measurement of other nuclear excited states in utN(0,σu2I)u_t\sim \mathcal N(0,\sigma_u^2 I)0Th” (Shigekawa et al., 2021) measures half-lives of several low-lying states, uses the Alaga rule to infer

utN(0,σu2I)u_t\sim \mathcal N(0,\sigma_u^2 I)1

and obtains the radiative half-life

utN(0,σu2I)u_t\sim \mathcal N(0,\sigma_u^2 I)2

The same paper states that this corresponds to a relative natural linewidth

utN(0,σu2I)u_t\sim \mathcal N(0,\sigma_u^2 I)3

which it describes as indicating superb performance for an optical nuclear clock. In that usage, the thorium TH-Score is naturally the combination of utN(0,σu2I)u_t\sim \mathcal N(0,\sigma_u^2 I)4, natural linewidth, and the feasibility of suppressing internal conversion so that the radiative transition becomes operationally accessible (Shigekawa et al., 2021).

These three thorium usages are related but non-identical. One emphasizes excitation rate, one spectroscopic inference of utN(0,σu2I)u_t\sim \mathcal N(0,\sigma_u^2 I)5, and one clock-linewidth performance. The common feature is that each compresses the value of utN(0,σu2I)u_t\sim \mathcal N(0,\sigma_u^2 I)6Th as a controllable nuclear-optical system into a small number of experimentally or theoretically central scalars (Porsev et al., 2010, Beloy, 2014, Shigekawa et al., 2021).

6. TTScore in speech synthesis and the problem of near-homonymous metrics

A near-homonymous but distinct metric family appears in “Objective Evaluation of Prosody and Intelligibility in Speech Synthesis via Conditional Prediction of Discrete Tokens” (Ulgen et al., 24 Sep 2025). The paper does not define TH-Score. It defines TTScore, with two components: utN(0,σu2I)u_t\sim \mathcal N(0,\sigma_u^2 I)7 for intelligibility via content tokens, and

utN(0,σu2I)u_t\sim \mathcal N(0,\sigma_u^2 I)8

for prosody via phoneme-level prosody tokens. Both are reference-free, text-conditioned, token-likelihood-based scores (Ulgen et al., 24 Sep 2025).

The distinction from TH-Score matters because the naming overlap can obscure the metric’s actual semantics. TTScore-int is a conditional content-likelihood score and TTScore-pro a conditional prosody-likelihood score; neither is threshold-based, threshold-free in the survival-analysis sense, nor related to TT-core ordering. The paper reports that TTScore-int and TTScore-pro achieve stronger correlations with human judgments than conventional intelligibility or prosody-focused baselines on SOMOS, VoiceMOS, and TTSArena, but these are properties of the TTScore family specifically (Ulgen et al., 24 Sep 2025).

This near-homonymy clarifies a broader point about TH-Score as an encyclopedia topic. The label is not a unified technical object but a recurrent score metaphor instantiated differently across fields. In system identification it names a rank-adaptive oracle-style trade-off; in evaluation theory it concerns threshold choice or threshold elimination; in tensor methods it is a weighted cut on an interaction graph; in thorium-229 physics it denotes performance figures for excitation, spectroscopy, and clocking; and in speech synthesis the relevant formalism is not TH-Score at all but TTScore (Zheng et al., 10 Jun 2025, Yuan et al., 2016, Kinjo et al., 29 Mar 2025, Shigekawa et al., 2021, Ulgen et al., 24 Sep 2025).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to TH-Score.