---
title: Almost Spanning Tree Measures
url: https://www.emergentmind.com/topics/almost-spanning-tree-measures
type: topic
---

# Almost Spanning Tree Measures

Searching arXiv for the cited papers and adjacent work on spanning-tree measures, distortion, and tree-likeness.
{"query":"Spanning Trees and Mahler Measure 1602.02797 arXiv"}
In the literature represented by these works, **almost spanning tree measures** can be understood as quantitative frameworks that evaluate how a graph, metric space, probability measure, or infinite combinatorial structure is controlled by spanning trees without restricting attention to the classical minimum-spanning-tree objective alone. The relevant notions include growth rates of spanning-tree complexity for periodic graphs, lightness–distortion tradeoffs for trees that are nearly minimum in weight, extremal compactness objectives over the family of spanning trees of a graph, deviation measures for metrics that are close to exact tree realizability, and probabilistic or measurable spanning-tree laws on infinite or random structures [1602.02797], [1609.08801], [2206.07104], [1505.06145], [2401.13238], [2509.15000], [2601.07740], [0912.4765], [2601.08396], [2404.06447].

## 1. Conceptual scope

Classical minimum spanning trees minimize total edge weight, but several of the cited works show that this criterion is often too narrow. In weighted graphs, the MST may have **average distortion** as bad as $\Omega(n)$, even though it is weight-optimal; in unweighted graphs, all spanning trees have the same number of edges and the same total edge weight $n-1$, so MST-style criteria do not distinguish them at all; and in finite metric spaces, the relevant question may be whether a fully labelled weighted tree on the same vertex set realizes the metric exactly rather than merely approximately [1609.08801], [2206.07104], [1505.06145].

This suggests a useful division of the subject into several regimes. One regime studies **exact tree control**, as in fully labelled tree realizability of a finite metric. A second studies **near-tree optimization**, where the tree is required to be almost minimum in weight while preserving distances much better than the MST. A third studies **extremal spanning-tree statistics**, such as average pairwise distance inside the tree. A fourth studies **measures on spanning trees** themselves, including uniform, minimal, or measurable spanning-tree laws on finite or infinite graphs. A fifth studies **outer optimization over the spanning-tree family**, where a non-tree problem, such as optimal transport, is reduced to a minimization over spanning trees. The phrase “almost spanning tree measures” is therefore best read as a family resemblance rather than as a single formal invariant.

A recurrent theme across these settings is that the spanning-tree family serves as a compressed but highly structured model class. Depending on context, the relevant statistic may be total weight, routing distortion, compactness, entropy growth, realizability error, isomorphism-type probability, end structure, or transport cost. The cited papers differ sharply in formalism, but they share the principle that spanning trees provide a canonical low-complexity scaffold against which more complicated combinatorial or geometric behavior can be measured.

## 2. Periodic complexity, Laplacian polynomials, and Mahler measure

For a finite connected graph, the **complexity** is the number of spanning trees, denoted $\tau(G)$. For a non-connected graph with connected components $G_1,\dots,G_\mu$, the multiplicative extension is
\[
T(G)=\tau(G_1)\cdots \tau(G_\mu).
\]
For an infinite graph $G$ with cofinite free $\mathbb{Z}^d$-symmetry, the relevant finite approximants are the finite-index quotients $G_\Lambda$, which may be disconnected, so the natural counting function is $T(G_\Lambda)$ rather than $\tau(G_\Lambda)$ [1602.02797].

The periodic structure yields a matrix Laplacian over
\[
\mathbb{Z}[x_1^{\pm1},\dots,x_d^{\pm1}],
\]
and the associated **Laplacian polynomial** is
\[
\Delta=\det(L),
\]
well defined up to multiplication by units in the Laurent polynomial ring. A combinatorial description is given by a Forman/Kenyon-type spanning forest formula:
\[
\Delta(G)=\sum_F \prod_{\text{Cycles of }F}(2-w-w^{-1}),
\]
where the sum is over cycle-rooted spanning forests and $w$ is the monodromy around each cycle. The logarithmic Mahler measure of $\Delta$,
\[
m(\Delta)=\int_0^1\cdots\int_0^1 \log\bigl|\Delta(e^{2\pi i\theta_1},\dots,e^{2\pi i\theta_d})\bigr|\,d\theta_1\cdots d\theta_d,
\]
is exactly the exponential growth rate of spanning-tree complexity in the quotients:
\[
\lim_{\langle \Lambda\rangle\to\infty}\frac{1}{|\mathbb{Z}^d/\Lambda|}\log T(G_\Lambda)=m(\Delta).
\]
Here $\langle \Lambda\rangle$ is the length of the shortest nonzero vector in $\Lambda$, so the limit is taken over fundamental domains growing in every direction [1602.02797].

The paper also gives an explicit product formula for $T(G_\Lambda)$:
\[
T(G_\Lambda)=\frac{1}{n_\Lambda}\prod_{\substack{(c_1,\dots,c_d)\in \Omega(\Lambda)\cap \mathbb{T}^d\\ (c_1,\dots,c_d)\neq 0}} \bigl|\Delta(c_1,\dots,c_d)\bigr|,
\]
with $n_\Lambda$ a normalization factor coming from the sizes of the connected components of $G_\Lambda$. The asymptotic theorem is then proved by interpreting this product as a Riemann sum converging to the Mahler integral. In this form, spanning-tree growth is identified with a robust algebraic invariant rather than with an ad hoc combinatorial asymptotic.

A central extremal statement is
\[
m(\Delta(G))\ge m(\Delta(\mathbb{G}_d)),
\]
where $\mathbb{G}_d$ is the $d$-dimensional grid graph. Since
\[
\Delta(\mathbb{G}_d)=2d-x_1-x_1^{-1}-\cdots-x_d-x_d^{-1},
\]
the grid minimizes complexity growth among periodic graphs with finitely many connected components. Moreover,
\[
m(\Delta(\mathbb{G}_d))\le \log(2d), \qquad
\lim_{d\to\infty}\bigl(m(\Delta(\mathbb{G}_d))-\log(2d)\bigr)=0,
\]
so the grid growth rate is asymptotic to $\log(2d)$. The paper also proves a gap theorem:
\[
m(\Delta(G))\ne 0 \implies m(\Delta(G))\ge \log 2.
\]
This lower bound is sharp, since doubling each edge of the $1$-dimensional grid yields Laplacian polynomial $2(2-x-x^{-1})$ with Mahler measure exactly $\log 2$ [1602.02797].

The same framework has a dynamical interpretation: $m(\Delta)$ is the topological entropy of the associated algebraic $\mathbb{Z}^d$-action on the Pontryagin dual of the coloring module. For $d\le 2$, the paper also connects the theory to determinant growth rates of alternating links arising from planar graphs via the medial construction, but the primary content is the equivalence
\[
\text{spanning-tree growth of }G \longleftrightarrow m(\Delta(G)).
\]
Within the present topic, this is one of the clearest examples of a spanning-tree measure that is not local, not purely combinatorial, and not reducible to a single finite tree count.

## 3. Almost minimum spanning trees and refined distortion measures

A different notion of “almost spanning tree” arises when the tree is required to remain nearly minimum in total weight while substantially improving distance preservation. For every weighted undirected graph and every parameter $0<\rho<1$, there exists a spanning tree $T$ with
\[
w(T)\le (1+\rho)\,w(\mathrm{MST}),
\]
and average distortion $O(1/\rho)$. This is described as an **almost minimum spanning tree**: a spanning tree whose total weight is only slightly above that of the MST, yet whose geometric quality is much better [1609.08801].

The paper formulates distortion through non-contractive embeddings $f:X\to Y$, meaning
\[
d_Y(f(x),f(y))\ge d_X(x,y)\qquad\forall x,y\in X.
\]
For graphs, the embedding is usually the identity into a subgraph, so the distortion of a pair is the stretch ratio
\[
\frac{d_H(u,v)}{d_G(u,v)}.
\]
The **average distortion** is the average over unordered pairs:
\[
\frac{1}{\binom{n}{2}}\sum_{\{u,v\}\in \binom{V}{2}} \frac{d_H(u,v)}{d_G(u,v)}.
\]
More generally, the paper defines $\ell_q$-distortion, with $\ell_1$ corresponding to average distortion and $\ell_\infty$ to worst-case distortion [1609.08801].

The more structural notions are **scaling distortion**, **coarse scaling distortion**, and **prioritized distortion**. Let
\[
R(v,\epsilon)=\min\{r:\ |B(v,r)|\ge \epsilon n\}.
\]
A point $u$ is $\epsilon$-far from $v$ if $d_X(u,v)\ge R(v,\epsilon)$. An embedding has scaling distortion $\gamma(\epsilon)$ if for every $\epsilon\in(0,1)$, at least a $(1-\epsilon)$-fraction of all pairs have distortion at most $\gamma(\epsilon)$. It has coarse scaling distortion $\gamma$ if every pair that is mutually $\epsilon/2$-far has distortion at most $\gamma(\epsilon)$. Given a ranking $\pi=v_1,\dots,v_n$, it has prioritized distortion $\alpha$ if for every $1\le j<i\le n$, the pair $(v_j,v_i)$ has distortion at most $\alpha(j)$. The paper proves that prioritized distortion and coarse scaling distortion are essentially equivalent up to reparameterization.

The main spanning-tree theorem is stated in scaling form: for any $0<\rho<1$, any graph contains a spanning tree with scaling distortion
\[
\tilde{O}\!\left(\sqrt{1/\epsilon}\right)/\rho
\]
and lightness $1+\rho$. Using the scaling-to-average-distortion lemma, this yields average distortion $O(1/\rho)$. The tradeoff is optimal in the sense that to get lightness $1+\rho$, average distortion must be $\Omega(1/\rho)$, even for spanning subgraphs rather than trees. Thus the asymptotically best relation is
\[
\text{lightness }1+\rho \quad\Longleftrightarrow\quad \text{average distortion }\Theta(1/\rho)
\]
[1609.08801].

The proof architecture is layered. First one builds a light spanner with prioritized distortion
\[
\tilde{O}(\log j)/\rho.
\]
The equivalence theorem converts this into coarse scaling distortion
\[
\tilde{O}(\log(1/\epsilon))/\rho.
\]
Then a tree-embedding theorem of [ABN15], as described in the paper, converts the spanner to a spanning tree with scaling distortion $O(1/\sqrt{\epsilon})$. A composition lemma for scaling distortions yields the final
\[
\tilde{O}(\sqrt{1/\epsilon})/\rho
\]
bound. A general lightness reduction further shows that if a spanner with lightness $\ell$ and distortion $t(u,v)$ exists for every weight function, then for every $0<\delta<1$ there exists a spanner of lightness
\[
1+\delta\ell
\]
and distortion
\[
t(u,v)/\delta.
\]
This reduction is what pushes lightness arbitrarily close to $1$ [1609.08801].

A common misconception is that MST optimality should already provide a useful routing tree. The paper gives the opposite picture: even on the unweighted cycle, every spanning tree stretches some adjacent pair by $\Theta(n)$, and the MST may have average distortion as bad as $\Omega(n)$. In this setting, the “almost” qualifier refers not to approximate connectivity but to near-optimal total weight coupled to much improved distance preservation.

## 4. Extremal compactness and spanning-tree-likeness of metrics

For unweighted connected graphs, one can optimize over the family of spanning trees using a distance statistic rather than weight. The paper on compact spanning trees defines
\[
C(G)=\frac{1}{n^2}\sum_{i=1}^n\sum_{j=1}^n D(v_i,v_j),
\]
where $D(v_i,v_j)$ is the hop distance. The **Most Compact Spanning Tree** and **Least Compact Spanning Tree** are
\[
T^*(G)=\arg\min_{T(G)\in\mathcal T(G)} C(T(G)),\qquad
T^\#(G)=\arg\max_{T(G)\in\mathcal T(G)} C(T(G)).
\]
Because every spanning tree of an unweighted graph has the same total edge weight $n-1$, compactness rather than weight distinguishes extremal trees [2206.07104].

The proposed algorithm is an iteratively greedy **rank-and-regress** elimination procedure. Starting from the full graph, it removes exactly one edge per iteration until only $n-1$ edges remain. The ranking is based on the matrix of relative forest accessibilities
\[
Q=(I+L)^{-1},
\]
the associated forest distance
\[
A_{ij}=Q_{ii}-Q_{ij}-Q_{ji}+Q_{jj},
\]
and the effective resistance distance
\[
R_{ij}=L^+_{ii}-L^+_{ij}-L^+_{ji}+L^+_{jj},
\]
where $L^+=\operatorname{pinv}(L)$. On an unweighted graph, if $R_{ij}=1$ on an edge $e_{ij}$, then the edge is a bridge and must not be deleted. Among non-bridge edges, the MCST deletes the edge with maximum forest-distance score, while the LCST deletes the minimum-score edge. The process converges after exactly $m-(n-1)$ deletions and stays connected throughout [2206.07104].

The empirical behavior matches graph-theoretic intuition. For complete graphs $K_n$, the method returns a star $S_n$ as the MCST and a path $P_n$ as the LCST. Tests on more than $500$ Erdős–Rényi graphs with $n\in[10,99]$ and $p\in[0.25,0.50]$, and about $360$ Barabási–Albert graphs with $n\in[10,99]$ and average degree between $2$ and $5$, show that the MCST consistently has lower average shortest-path distance than random spanning trees, while the LCST has higher distance. The stated running time is
\[
O\!\left(m(2n^3+m)\right)=O(mn^3),
\]
with worst-case $O(n^5)$ when $m=O(n^2)$ and about $O(n^4)$ for sparse graphs where $m=O(n)$ [2206.07104].

A related but conceptually distinct line of work studies whether a finite metric is itself exactly or approximately tree-realizable on the same vertex set. Let $M=(X,d_M)$ be a finite metric space. A fully labelled positive-weighted tree $T$ on the same vertex set realizes $M$ if
\[
d_T(x,x')=d_M(x,x')\qquad\forall x,x'\in X.
\]
If such a spanning-tree representation exists, then the basic geodesic graph $G_M$ is the **unique minimum spanning tree** in the complete weighted graph $K_M$, and it is the unique fully labelled tree on $X$ realizing $M$ [1505.06145].

The governing criterion is the **fourth-point condition**: for every triple $x,y,z\in X$, there exists $p^*\in X$ such that
\[
d_M(x,p^*)+d_M(y,p^*)+d_M(z,p^*)
=\frac12\bigl(d_M(x,y)+d_M(y,z)+d_M(z,x)\bigr).
\]
Under the **tie-breaking rule**
\[
d_M(x,y)\neq d_M(u,v)\quad \text{for all distinct pairs } \{x,y\}\neq \{u,v\},
\]
the paper proves that $M$ is a spanning tree metric space if and only if it satisfies the fourth-point condition. The fourth point is unique whenever it exists, and it is the median of $\{x,y,z\}$ [1505.06145].

The approximate deviation from exact tree behavior is measured by the **roundaboutness**
\[
\rho = \max_{x,y,z\in X}\min_{i\in X}
\frac{d_M(x,i)+d_M(y,i)+d_M(z,i)}
{d_M(x,y)+d_M(y,z)+d_M(z,x)}-\frac12.
\]
Here $\rho=0$ means exact spanning-tree realizability under tie-breaking, while $\rho>0$ quantifies deviation from exact tree-likeness. The paper interprets $\rho$ as the goodness-of-fit of the MST to $M$. It also gives a path analogue via the three-point condition and states algorithmic complexities $O(n^4)$ for spanning-tree-likeness and $O(n^3)$ for spanning-path-likeness [1505.06145].

A notable caution is that the fourth-point condition alone is not sufficient without tie-breaking: complete graphs with uniform edge length can satisfy it yet fail to be trees on the same vertex set. This clarifies that “tree-like” in this work means fully labelled spanning-tree-like, not merely classical Buneman tree-likeness.

## 5. Directed, measurable, and random spanning-tree measures

Spanning-tree measures become more delicate on infinite or random structures. In the directed setting, the **minimal spanning arborescence** is the analogue of the MST for a graph in which each undirected edge has both orientations. A spanning arborescence with boundary $\partial$ requires every vertex $v\neq\partial$ to have exactly one outgoing edge, $\partial$ to have zero outgoing edges, and no directed cycles. For i.i.d. continuous weights on oriented edges, the MSA is the unique spanning arborescence minimizing total weight. Unlike the undirected MST, however, the law of the MSA depends on the weight distribution; the paper gives an explicit counterexample showing different probabilities under Exponential$(1)$ and Uniform$[0,1]$ weights. Standard Kruskal and Prim methods do not apply. Instead, the paper develops the Chu–Liu/Edmonds/Bock recursion and the associated **loop contracting random walk**, which is described as similar to loop-erased random walk except that loops are contracted rather than erased [2401.13238].

For infinite graphs, the paper defines the wired boundary MSA limit through an exhaustion. It proves almost-sure existence of the limit on bounded subdivisions of transient trees with no degree-$2$ vertices under i.i.d. Exponential$(1)$ weights, and proves that for invariantly nonamenable unimodular random rooted graphs satisfying a transient CLEB-walk hypothesis, the wired MSA limit exists almost surely, has infinitely many infinite components, and each component is one-ended. The paper also shows that the CLEB walk is a limit of Wilson’s algorithm as $\beta\to\infty$ in a weighted uniform spanning arborescence model with conductances $e^{-\beta U_{\vec e}}$ [2401.13238].

In measurable graph combinatorics, the relevant object is a **measurable one-ended spanning tree**. For a locally finite, one-ended, mcp graph $\mathcal G$ on a standard probability space $(X,\nu)$, the main theorem states that $\mathcal G$ is hyperfinite if and only if it admits a measurable one-ended spanning tree almost everywhere. Here hyperfiniteness means that $\mathcal G$ can be written as an increasing union of measurable graphs with finite connected components almost everywhere. The proof uses substantial subgraphs, the core of a spanning tree, boundary connectivity arguments, contraction of finite non-core pieces, and a final deletion step based on $3$-edge-connectivity and Nash–Williams tree packing. In this setting, hyperfiniteness is identified as the exact obstruction to measurable one-ended treeings [2509.15000].

For the **uniform spanning tree** on $\mathbb{Z}^2$, the spanning-tree measure itself has a precise fractal geometry. The paper proves that the UST is one-sided and that simple random walk on it has exponents
\[
d_f=\frac85,\qquad d_w=\frac{13}{5},\qquad d_s=\frac{16}{13},
\]
where $d_f$ is the volume-growth dimension, $d_w$ the walk dimension, and $d_s$ the spectral dimension. Writing $G(n)$ for the expected length of an infinite loop-erased random walk until it exits the Euclidean ball of radius $n$, the paper uses
\[
G(n)\approx n^{5/4},\qquad g(t)\approx t^{4/5}
\]
for the inverse $g$, and proves
\[
p_{2n}^\omega(0,0)\approx n^{-8/13},\qquad
\tau_R\approx R^{13/5},\qquad
\tau_r\approx r^{13/4}.
\]
Thus the random spanning-tree measure induces a random fractal metric space whose walk and spectral behavior are tree-controlled but far from Euclidean [0912.4765].

A different distributional phenomenon appears in dense finite graphs. For connected $(1\pm 10^{-10})d$-regular graphs with sufficiently large degree, the uniformly random spanning tree $\mathcal T$ is strongly anticoncentrated among isomorphism types:
\[
\Pr[\mathcal T\simeq_{\mathrm{iso}} T]\le e^{-n/2000}
\]
for every fixed $n$-vertex tree $T$. Consequently, the number of unlabeled spanning-tree isomorphism classes satisfies
\[
|T_{\mathrm{unlabeled}}(G)|\ge e^{n/2000}.
\]
The proof introduces a graph-theoretic balls-into-bins model on almost regular bipartite graphs, establishes Poisson-like occupancy behavior using Talagrand concentration, and combines this with a counting theorem for labeled tree embeddings and a lower bound on the total number of spanning trees [2601.07740].

Taken together, these results show that spanning-tree measures can encode very different structural phenomena: directed optimization, measurable one-endedness, fractal random geometry, and anticoncentration over isomorphism classes. A common misconception is that “spanning-tree measure” always means the uniform measure on finite spanning trees. In fact, the cited works use the term across optimization, probability, and measurable combinatorics in substantially different senses.

## 6. Spanning-tree optimization beyond graph connectivity: transport and centrality

In discrete optimal transport on a finite metric space whose ground cost is the shortest-path metric of a weighted connected graph $G=(\mathcal X,E_G,w)$, the Kantorovich distance can be written as an outer minimization over rooted spanning trees:
\[
K_{d_G}(\mu,\nu)=\min_{T\in ST}K_{d_T}(\mu,\nu)
=\min_{T\in ST}\sum_{x\in\mathcal X\setminus\{r\}} d_T(x,x^+)\,|\mu(T_x)-\nu(T_x)|.
\]
Here $x^+$ is the parent of $x$ and $T_x$ is the subtree rooted at $x$. The root is only a notational device; the value is invariant under rerooting. Thus the transport problem on $G$ is reduced to choosing an optimal spanning tree for the pair $(\mu,\nu)$ [2601.08396].

For a rooted weighted tree $T$, defining
\[
\Xi_T(x)=\mu(T_x)-\nu(T_x),
\]
the tree transport cost is
\[
K_{d_T}(\mu,\nu)=\sum_{x\neq r} d_T(x,x^+)\,|\Xi_T(x)|.
\]
The corresponding Kantorovich potential is
\[
u_T(y)=\sum_{y\lesssim x\neq r} d_T(x,x^+)\,\operatorname{sgn}(\Xi_T(x)),\qquad u_T(r)=0,
\]
and under **weak non-degeneracy**
\[
\mu(X')\neq \nu(X')\qquad \forall\,\emptyset\neq X'\subsetneq \mathcal X,
\]
this potential is unique up to an additive constant. The paper also gives a dynamic-programming procedure that constructs an optimal coupling for $d_T$ by leaf peeling once an optimal tree is known, and a simulated-annealing algorithm on the space of spanning trees to solve the outer minimization [2601.08396].

A different outer optimization over spanning trees appears in Euclidean data analysis. The **central spanning tree** is defined on a complete Euclidean graph with terminal coordinates $X_V=\{x_1,\dots,x_{|V|}\}\subset\mathbb{R}^n$ by
\[
\operatorname{CST}
=\underset{\tree}{\arg\min}
\sum_{(i,j)\in E_{\tree}}
\bigl(m_{ij}(1-m_{ij})\bigr)^\alpha\,\|x_i-x_j\|,
\]
where $m_{ij}$ is the normalized size of one component produced by deleting edge $(i,j)$. Since $m_e(1-m_e)$ is proportional to edge betweenness centrality in a tree, the objective weights edge length by edge centrality. The **branched central spanning tree** adds Steiner points and optimizes both topology and branch-point positions [2404.06447].

The parameter $\alpha\in[0,1]$ interpolates among classical tree objectives. At $\alpha=0$, the CST becomes the MST and the BCST becomes the Steiner tree. At $\alpha=1$, the CST becomes proportional to the minimum routing cost tree / optimum distance spanning tree. The paper states that as $\alpha$ increases, the tree becomes more robust to perturbations but increasingly star-like; for $\alpha\to\infty$, or for $\alpha>1$ and large $N$, the optimum tends to a star, while for $\alpha\to-\infty$ it tends to a path graph. Both CST and BCST are NP-hard, so the paper proposes the heuristic **mSTreg**, which alternates between topology updates and geometric optimization of Steiner points by IRLS [2404.06447].

These two papers show complementary roles for spanning trees outside traditional graph algorithms. In optimal transport, the tree is an exact optimizer for a reformulated Kantorovich problem. In Euclidean data analysis, the tree is a tunable skeleton balancing local fidelity against global robustness. In both cases, the objective is not “find a spanning tree” per se, but rather “express a more complex optimization problem through a spanning-tree search space.”

The broader implication is that almost spanning tree measures are not confined to counting trees or choosing the cheapest one. They include asymptotic complexity invariants, distortion profiles, compactness extrema, realizability errors, measurable existence criteria, probabilistic laws on random trees, and variational reductions over the spanning-tree family. What unifies them is the use of spanning trees as a structurally rigid yet sufficiently expressive class against which combinatorial, metric, probabilistic, and analytic behavior can be measured.

Source: https://www.emergentmind.com/topics/almost-spanning-tree-measures