Expression-Distance Metric Overview
- Expression-Distance Metric is a design principle where explicit algebraic expressions of base distances encode similarity, dependence, and classification confidence.
- It transforms conventional dissimilarities into valid metrics using methods like angular transforms and normalization, enabling robust indexing, clustering, and embedding.
- The approach enhances interpretability and invariance across applications—from metric learning and probability measures to symbolic transformations and set comparisons.
Searching arXiv for the supplied topics and papers to ground the article in current records. Across the cited literature, “Expression-Distance Metric” can be interpreted as a family of constructions in which similarity, dependence, or classification confidence is written explicitly as a function of distances, ratios of distances, or expectations of distances, rather than through unconstrained logits or purely coordinate-based summaries. In this sense, the term covers correlation-derived metrics for expression profiles, distance-ratio objectives in metric learning, metrics on sets and probability measures built from pairwise distances, point-to-set distances such as Scott distance, and lifted distances on symbolic transformations such as finite-state transductions (Kim et al., 2022, Lyons, 2011, Aiswarya et al., 2024).
1. Conceptual scope
The unifying feature is that the primary object is an explicit distance expression. In metric learning, this appears as a probability model written directly in terms of distances to class representatives; in dependence testing, as expectations of centered distance kernels; in set and distribution comparison, as averages or norms built from pairwise distances; and in transduction comparison, as the supremum of output distances over all inputs (Kim et al., 2022, Lyons, 2011, Fujita, 2011, Aiswarya et al., 2024).
This suggests a broad but technically coherent interpretation: an expression-distance metric is not a single canonical metric, but a design principle in which the comparison rule is itself an algebraic expression over a base metric or similarity. Some constructions yield genuine metrics; others yield metric-like quantities, losses, or probabilistic scores whose geometric meaning is explicit.
| Setting | Distance expression | Key property |
|---|---|---|
| Correlation profiles | , | True metrics |
| Metric learning | Scale-invariant probabilities | |
| Finite sets | from averaged pairwise distances | Metric on non-empty finite subsets |
| Probability measures | Genuine metric in strong-negative-type spaces | |
| PDFs in | Metric with Gaussian-mixture closed form | |
| Transductions | Metric on partial functions |
A recurring consequence is interpretability of geometry. Distances become the primary quantities, and relative position, scaling behavior, or transport structure can often be read directly from the defining expression.
2. Correlation-derived metrics for expression profiles
In expression analysis, a profile is represented as a vector , and the Pearson correlation coefficient is
A central distinction in this literature is between a dissimilarity and a metric. The widely used Pearson dissimilarity 0 is not a metric, because the triangle inequality can fail. The same is true for the sign-invariant form 1. By contrast, the paper proves that
2
is a metric, and that the sign-invariant
3
is also a metric (Solo, 2019).
A second line of work systematizes these constructions through angular distance and metric-preserving transforms. For cosine similarity, Pearson correlation, or Spearman correlation 4, the angular distance
5
is a metric. Applying increasing, concave transforms yields further metrics. This produces two classes. In the first, anti-correlated objects are maximally far apart, with examples
6
In the second, correlated and anti-correlated objects are collated, with examples
7
The same formulas apply to Spearman correlation by replacing 8 with 9 (Dongen et al., 2012).
For centered data, Pearson correlation is the cosine similarity of centered vectors, so the sine transform yields the metric
0
This is important in practice because many clustering and embedding procedures implicitly assume metric structure. The literature therefore separates correlation-based quantities that are merely convenient dissimilarities from those that satisfy the triangle inequality and can support metric-based indexing, clustering, and embedding (Solo, 2019, Dongen et al., 2012).
3. Learned distance formulations in embedding spaces
In metric learning, the “expression-distance” idea appears when probabilities and losses are written directly in terms of distances to class representatives. In the prototypical-network setting, if 1 is the prototype of class 2 and 3, the distance-ratio formulation defines
4
with learnable 5, and optimizes cross-entropy over query points. This formulation is exactly invariant to global scaling of the embedding, because 6 cancels in the ratio. It also has an optimal-confidence property: if prototypes are pairwise separated, then 7 as 8, and 9 as 0 for 1. On CUB and mini-ImageNet, the paper reports that this formulation generally enables faster and more stable metric learning than the softmax-based formulation, with improved or comparable generalization performance (Kim et al., 2022).
A second formulation keeps the metric explicit in the input space. In generalized Mahalanobis metric learning,
2
For a triplet 3, the paper defines the adversarial margin as the Euclidean distance from 4 to the nearest decision boundary between the target neighbor 5 and the impostor 6. Its squared closed form is
7
This yields a perturbation loss that can be added to LMNN, SCML, or proxy-based methods, and the resulting enlarged adversarial margin is used to derive certified robustness and generalization guarantees via algorithmic robustness (Yang et al., 2020).
Other learned expression-distance metrics make the structural form explicit. Boosted sparse non-linear distance metric learning uses
8
with 9 built as a sum of sparse rank-one updates, so that the final metric is positive semidefinite, low rank, and element-wise sparse. Nonlinearity is introduced through hierarchical expansion of interactions in the feature map (Ma et al., 2015). A different hybrid method combines a weighted Euclidean term with a class-mediated latent term: 0 where 1 is a soft class posterior and 2 models missing-feature effects conditioned on class. This construction is designed for settings where similarity ratings and class labels both inform perceived distance (Kao et al., 2012).
4. Sets, measures, and distributions
One important strand of the literature defines distances between complex objects by averaging the base metric over their components. For non-empty finite subsets 3 of a metric space 4, the group-average distance
5
is symmetric and satisfies the triangle inequality, but is not a metric because it does not satisfy identity. The corrected quantity
6
is a metric on the collection of all non-empty finite subsets. Under the discrete metric, it reduces to the Jaccard distance, and the construction extends hierarchically to sets of sets and, by integration, to measurable and fuzzy sets (Fujita, 2011).
For probability measures and dependence structure, the central object is again an expression of distances. In a metric space 7, with 8 a probability measure of finite first moment,
9
and the centered kernel is
0
Distance covariance is then
1
with empirical versions obtained by double-centering distance matrices. In spaces of strong negative type, 2 if and only if 3 and 4 are independent, and
5
is a genuine metric on probability measures with finite first moments (Lyons, 2011).
Optimal transport on finite metric spaces provides another explicit expression. For 6 on a finite metric space 7, the Kantorovich distance equals the Kantorovich–Bernstein or Arens–Eells norm of 8. On a weighted tree rooted at 9,
0
and for a general weighted graph it is the minimum of the corresponding tree distances over all spanning trees (Montrucchio et al., 2019).
For continuous probability densities in 1, the proposed probabilistic distance metric is built from the normalized cross-information potential
2
yielding
3
This is a true metric on square-integrable PDFs. For Gaussian mixtures 4 and 5, the metric has a closed-form expression in terms of pairwise Gaussian product integrals, which enables optimization-based Gaussian mixture reduction (Sajedi et al., 2023).
Finite metric-measure spaces admit yet another higher-order construction. For a cloud 6 and an 7-point simplex 8, the Relative Distance Distribution 9 packages the internal distance matrix 0 and the lexicographically ordered matrix 1 of distances from 2 to the remaining points, modulo permutations of the simplex. The Simplexwise Distance Distribution
3
is an isometry invariant. Metrics on SDDs are then built via a metric 4 on RDDs together with Linear Assignment Cost or Earth Mover’s Distance. These metrics are Lipschitz continuous under point perturbations, and order 5 already distinguishes examples that share the same pairwise or pointwise distance distributions (Kurlin, 2023).
5. Point-to-set distances, generalized metrics, and symbolic transformations
Expression-distance constructions also arise when the second argument is a subset or a transformation rather than another point. On a generalized metric space 6, a Scott weight is a weight 7 that respects forward Cauchy nets and Yoneda limits. The Scott distance is
8
It turns 9 into an approach space and recovers the original metric on singletons: 0 Its topological coreflection, the c-Scott topology, lies between the 1-Scott topology and the generalized Scott topology, and every injective 2 approach space is a cocomplete and continuous metric space equipped with its Scott distance (Li et al., 2016).
The generalized-metrics literature relaxes Fréchet’s axioms in other directions. Partial metrics and strong partial metrics allow non-zero self-distance and even negative values, while partial 3-4etrics and strong partial 5-6etrics compare 7-tuples rather than pairs. A partial metric 8 satisfies
9
and induces the classical metric
0
The thesis further shows that a DNA alignment scoring function can be modeled as a strong partial metric under explicit conditions on match, mismatch, and gap penalties, allowing topological and fixed-point methods to be applied to a comparative function that is not a classical metric (Assaf, 2016).
For symbolic transformations, a word metric can be lifted to a metric on partial transductions. If 1 are partial functions and 2 is a metric on words, then
3
This is a metric on partial word-to-word functions. For common integer-valued edit distances, including Hamming, transposition, conjugacy, and Levenshtein-family distances, closeness and 4-closeness are decidable for functional transducers, so the distance is computable. The same quantity is equivalent to the diameter of a rational relation, and both are specific instances of the index problem of rational relations (Aiswarya et al., 2024).
6. Recurring principles, misconceptions, and limitations
A first recurring principle is that not every useful dissimilarity is a metric. The expression-analysis literature gives the clearest examples: 5 and 6 fail the triangle inequality, whereas 7 and 8 satisfy it (Solo, 2019). The metric-preserving-function framework sharpens this point: increasing concave transforms preserve metricity, while strictly convex transforms such as 9 do not in general (Dongen et al., 2012).
A second principle is invariance. The distance-ratio formulation is exactly invariant to global rescaling of embeddings, so its loss depends on relative distances rather than norm inflation (Kim et al., 2022). SDD-based metrics are Lipschitz continuous under perturbations of the underlying points (Kurlin, 2023). The probabilistic metric 00 is bounded in 01, which stabilizes comparison of densities (Sajedi et al., 2023). These constructions suggest that explicit distance expressions are often chosen precisely to control undesirable dependence on scaling, ordering, or noise.
A third principle is that characterization theorems depend on geometric hypotheses. Distance covariance characterizes independence if and only if the underlying metric spaces are of strong negative type (Lyons, 2011). Scott distance captures approach-theoretic structure because Scott weights are tied to Yoneda continuity (Li et al., 2016). In finite optimal transport, explicit formulas are available on trees and extend to general graphs through minimization over spanning trees (Montrucchio et al., 2019). These are not merely computational conveniences; they delimit where the distance expression has full structural meaning.
The main limitations are likewise recurrent. Higher-order invariants such as 02 become combinatorially more expensive as 03 grows, even though they gain discriminative power (Kurlin, 2023). In transduction spaces, the bounded index problem for rational relations is undecidable in general, even though it becomes decidable for the edit-distance families treated in the paper (Aiswarya et al., 2024). Metric learning formulations that are geometrically cleaner may introduce extra parameters or representative-choice issues, as in the learnable 04 of the distance-ratio model or the representative dependence of multi-shot few-shot learning (Kim et al., 2022). This suggests that the central trade-off in expression-distance design is between explicit geometry, metric validity, invariance, and tractability.
Taken together, these works present “Expression-Distance Metric” less as a single named object than as a recurring mathematical strategy: define similarity, dependence, or transformation discrepancy directly from distances native to the domain, and then study which expressions preserve metric axioms, which yield useful invariances, and which remain computable at the scale of the intended application.