---
title: Sample-Supported Expressibility Law
url: https://www.emergentmind.com/topics/sample-supported-expressibility-law
type: topic
---

# Sample-Supported Expressibility Law

Searching arXiv for the specified paper and closely related work on expressibility and variational quantum ansatzes.
arXiv search: 2606.12211
The sample-supported expressibility law is an information-theoretic principle for circuit-based quantum learning that links the expressibility of a quantum ansatz to the number of copies of an unknown quantum state available for learning. In the formulation developed in "Quantum Occam Learning: Sample-Supported Expressibility for Circuit-Based Quantum Learning" [2606.12211], the law states that, for $n$-qubit pure states preparable with at most $G$ two-qubit gates, trace-distance accuracy $\varepsilon$ can be supported by $M$ copies only up to a gate budget $G_{\rm supported}(M,\varepsilon)=\widetilde\Theta(\min\{2^n,\;M\varepsilon^2\})$ in the circuit-limited regime. The same phrase, "sample-supported expressibility law," also appears in work on variational quantum ansatzes where expressibility is defined through covering numbers of a hypothesis space and bounded above and below in terms of the number of trainable gates [2311.01330]. Across these usages, the common theme is that expressibility is not an absolute architectural property alone; it is meaningful only relative to finite statistical support.

## 1. Formal setting and core definitions

In the circuit-based learning framework of [2606.12211], the ambient Hilbert-space dimension is $N=2^n$. The trace distance between states $\hat\rho$ and $\hat\sigma$ is
$$
D(\hat\rho,\hat\sigma)=\tfrac12\|\hat\rho-\hat\sigma\|_1.
$$
The circuit-generated class is defined by
$$
U_{n,G}=\{\text{$n$-qubit unitaries implementable with $\le G$ two-qubit gates}\},
$$
and
$$
S_{n,G}=\{\hat\sigma_U=U|0^n\rangle\langle0^n|U^\dagger:U\in U_{n,G}\}.
$$
For an arbitrary source $\hat\rho$, the best $G$-gate approximation error is
$$
d_G(\hat\rho)=\inf_{\hat\sigma\in S_{n,G}}D(\hat\rho,\hat\sigma),
$$
and the approximate circuit complexity at tolerance $\eta$ is
$$
C_\eta(\hat\rho)=\min\{G:d_G(\hat\rho)\le \eta\}.
$$
The covering number $N(S_{n,G},D,\varepsilon)$ is the size of the smallest $\varepsilon$-net in trace distance, and the sample complexity $M$ is the number of independent copies of $\hat\rho$ available for arbitrary collective POVMs [2606.12211].

A key estimate, stated as Lemma 1 in [2606.12211], bounds the metric entropy:
$$
\log N(S_{n,G},D,\varepsilon)\le C_{\rm ent}\,G\,\log\!\bigl(C_{\rm arch}\,nG/\varepsilon\bigr),
$$
for $0<\varepsilon<1/2$. In $\widetilde O$ notation, this becomes $\log N(S_{n,G},D,\varepsilon)=\widetilde O(G)$. This estimate is the structural basis for the ensuing Occam-style sample laws.

A distinct but related definition appears in the variational-quantum setting of [2311.01330]. There, for an $N$-qudit parameterized circuit $U(\theta)$ built from $N_{nt}$ trainable gates, the hypothesis space is
$$
H=\{\mathrm{Tr}[U(\theta)^\dagger O U(\theta)\rho]\;|\;\theta\in\Theta\},
$$
and the $\varepsilon$-covering number $C(\varepsilon)\equiv\mathcal N(H,\varepsilon,|\cdot|)$ is the smallest cardinality of a set of samples that $\varepsilon$-covers the whole hypothesis space. In that formulation, a large $C(\varepsilon)$ means high expressibility, whereas a small $C(\varepsilon)$ signals low expressibility [2311.01330].

## 2. Statement of the law

For fixed target trace-distance accuracy $\varepsilon$ and confidence $1-\delta$, [2606.12211] states that in the circuit-limited regime, namely $G\ll 2^n$, the largest gate count for which one can uniformly guarantee trace-distance error $\le \varepsilon$ from only $M$ copies satisfies, up to logarithmic factors,
$$
G_{\rm supported}(M,\varepsilon)=\widetilde\Theta\!\bigl(\min\{2^n,\;M\varepsilon^2\}\bigr).
$$
Equivalently, to have a uniform guarantee at accuracy $\varepsilon$, one needs
$$
\sqrt{\tfrac{G}{M}}\lesssim \varepsilon
\qquad\Longleftrightarrow\qquad
G\lesssim M\varepsilon^2,
$$
up to logarithmic factors, until the pure-state tomography barrier $G\sim 2^n$ is reached [2606.12211].

The same paper summarizes the law in four clauses. In the realizable setting, $M=\widetilde\Theta(G/\varepsilon^2)$ for $S_{n,G}$. In the agnostic setting, one can achieve
$$
D\le d_G(\hat\rho)+\widetilde O(\sqrt{G/M}).
$$
In the adaptive setting, one competes with
$$
\inf_G\{d_G(\hat\rho)+\widetilde O(\sqrt{G/M})\}.
$$
Taken together, these yield the sample-supported expressibility condition
$$
G\lesssim \min\{2^n,\;M\varepsilon^2\}\quad(\text{up to logs}),
$$
which treats learnable circuit complexity as constrained by sample support rather than fixed a priori [2606.12211].

In [2311.01330], the term "sample-supported expressibility law" refers instead to a two-sided covering-number bound for variational ansatzes:
$$
\left(\frac{3N_{nt}\|O\|}{8\varepsilon}\right)^{d^{2k}N_{nt}}
\le C(\varepsilon)\le
\left(\frac{7N_{nt}\|O\|}{\varepsilon}\right)^{d^{2k}N_{nt}}.
$$
Here $\varepsilon$ is the desired worst-case approximation error in operator outputs, $N_{nt}$ is the total number of trainable gates, $k$ is the maximum qudit-arity, $d$ is the local dimension, and $\|O\|$ is the operator norm of the measured observable [2311.01330]. This usage is not identical to the Occam-theoretic law of [2606.12211], but both are organized around covering numbers and architectural capacity.

## 3. Realizable and agnostic Occam bounds

The upper-bound argument in [2606.12211] proceeds in three steps. First, Lemma 1 provides an $\varepsilon$-net $\mathcal H\subset S_{n,G}$ with $\log|\mathcal H|=\widetilde O(G)$. Second, a finite quantum hypothesis-selection result, stated as Lemma 2, gives an information-theoretic procedure which, for a finite family $\mathcal H$ and any unknown state $\hat\rho$, uses
$$
M\ge \frac{C}{\varepsilon^2}\Bigl[\log|\mathcal H|+\log(1/\delta)\Bigr]
$$
copies and, with probability at least $1-\delta$, outputs $\hat\sigma\in\mathcal H$ such that
$$
D(\hat\rho,\hat\sigma)\le c\,\inf_{\tau\in\mathcal H}D(\hat\rho,\tau)+\varepsilon.
$$
Third, in the realizable case $\hat\rho\in S_{n,G}$, the best net point has distance at most $\varepsilon$, which yields
$$
M_{\rm real}(n,G,\varepsilon,\delta)\le
\frac{C}{\varepsilon^2}\Bigl[G\log(nG/\varepsilon)+\log(1/\delta)\Bigr]
=\widetilde O\!\bigl(\tfrac{G}{\varepsilon^2}\bigr)
$$
[2606.12211].

The agnostic extension introduces the best $G$-gate approximation error $d_G(\hat\rho)$ and proves that with $M$ copies one can learn up to the best $G$-gate approximation error plus a statistical penalty $\widetilde O(\sqrt{G/M})$ [2606.12211]. In the notation used in the exposition, the approximation-estimation trade-off is
$$
\Phi_M(G;\hat\rho)=d_G(\hat\rho)+O\!\Bigl(\sqrt{\tfrac{G}{M}}\Bigr).
$$
This expresses the central operational meaning of the law: adding gates can reduce approximation error, but only at the cost of a larger statistical term.

A plausible implication is that expressibility in this framework is not identified with the sheer size of a hypothesis class. Rather, expressibility is meaningful only to the extent that finite data can discriminate among the hypotheses that the circuit class makes available. That interpretation is directly aligned with the paper’s formulation that expressibility is statistically meaningful only insofar as it can be learned from finitely many copies of an unknown quantum state [2606.12211].

## 4. Lower bounds, packing, and tomography saturation

The matching lower bound in [2606.12211] is obtained through a packing argument. For $G\ll 2^n$, one can embed an exponentially large packing $\mathcal P\subset S_{n,G}$ with
$$
\log|\mathcal P|\ge c\,G
$$
and pairwise trace distance at least $2\varepsilon$. Any learner that with $M$ copies achieves uniform error $\le \varepsilon$ must distinguish these equiprobable states. By Fano’s inequality, the transcript must convey $\Omega(\log|\mathcal P|)$ bits, while the Holevo bound implies that each copy can carry only $O(\varepsilon^2)$ bits of information about a packing separated by $\varepsilon$. Consequently,
$$
M\times O(\varepsilon^2)\gtrsim G
\qquad\Longrightarrow\qquad
M\gtrsim G/\varepsilon^2.
$$
Equivalently, uniform learning at error $\varepsilon$ requires $G\lesssim M\varepsilon^2$ [2606.12211].

This lower bound matches the upper bound up to logarithmic factors and is therefore not merely a sufficient-condition statement. The law is presented as a genuine threshold principle: at trace-distance accuracy $\varepsilon$, $M$ samples can support only $G_{\rm supported}\simeq M\varepsilon^2$ gates, up to logarithmic factors and the saturation imposed by tomography [2606.12211].

Tomography saturation enters once $G$ exceeds the threshold for universal pure-state synthesis, $G\sim 2^n$. At that point the class $S_{n,G}$ stops growing in metric entropy, and the lower bound becomes
$$
M\gtrsim \frac{2^n}{\varepsilon^2},
$$
which is the usual tomography scale. Thus the full law is
$$
G_{\rm supported}(M,\varepsilon)=\widetilde\Theta\!\bigl(\min\{2^n,\;M\varepsilon^2\}\bigr)
$$
[2606.12211].

A common misunderstanding is to treat the tomography barrier and the circuit-limited regime as unrelated phenomena. The formulation in [2606.12211] explicitly connects them: the linear-in-$M\varepsilon^2$ gate-support relation holds only until the pure-state tomography barrier is reached, after which increasing $G$ does not enlarge the relevant metric entropy.

## 5. Adaptive model selection and circuit complexity as a statistical resource

A central contribution of [2606.12211] is to remove the need to know $G$ in advance. The adaptive model-selection theorem uses a nested hierarchy $\{S_{n,G_j}\}$ and a penalized tournament, described as structural risk minimization, to select a hypothesis $\hat\sigma$ satisfying with high probability
$$
D(\hat\rho,\hat\sigma)\le
C\inf_G\Bigl[d_G(\hat\rho)+\widetilde O\!\bigl(\sqrt{G/M}\bigr)\Bigr].
$$
The theorem is summarized as selecting the circuit complexity justified by the data and establishing an oracle inequality [2606.12211].

This result recasts bounded circuit complexity as a model-selection principle for quantum machine learning [2606.12211]. Rather than treating $G$ as a fixed promise supplied externally, the framework makes circuit complexity adaptive: the data justify as many gates as make the reduction in approximation error worth the increase in the statistical penalty $\sqrt{G/M}$.

If one insists on a uniform trace-distance guarantee at level $\varepsilon$, the penalty term forces
$$
\sqrt{G/M}\le \varepsilon,
$$
hence $G\le M\varepsilon^2$ [2606.12211]. Gates above that scale are described as lying in the unsupported regime, with too many distinguishable hypotheses compared to the available samples. In this sense, circuit complexity becomes an adaptive statistical resource rather than a static architectural promise.

A plausible implication is that the law provides a principled alternative to ansatz selection based solely on hardware constraints or heuristic notions of expressibility. Within this framework, the appropriate circuit size is determined by the approximation-estimation trade-off induced by the copy budget.

## 6. Relation to variational-ansatz expressibility

The variational-quantum literature uses "expressibility" in a different but related sense. In [2311.01330], expressibility is defined as the covering number of the hypothesis space associated with an observable $O$ and input state $\rho$. Prior work by Du et al. is cited there as establishing the upper bound
$$
C(\varepsilon)\le
\left(\frac{7N_{nt}\|O\|}{\varepsilon}\right)^{d^{2k}N_{nt}}
$$
for $0<\varepsilon\le 1/10$, using the operator norm on the unitary group. The paper then derives a matching lower bound by a gate-by-gate covering argument:
$$
\mathcal N(H_{\rm circ},2\varepsilon,\|\cdot\|)\ge
\left(\frac{3N_{nt}\|O\|}{8\varepsilon}\right)^{d^{2k}N_{nt}},
$$
and, via a bi-Lipschitz trace map with unit constants, obtains the same lower bound for $C(\varepsilon)$ [2311.01330].

Taking logarithms gives the two-sided inequality
$$
d^{2k}N_{nt}\cdot \log\!\left(\frac{3N_{nt}\|O\|}{8\varepsilon}\right)
\le \log C(\varepsilon)\le
d^{2k}N_{nt}\cdot \log\!\left(\frac{7N_{nt}\|O\|}{\varepsilon}\right)
$$
[2311.01330]. The assumptions include use of the operator norm for unitary-group coverings, the induced trace-distance norm on the hypothesis space, uniform sampling over the continuous parameter space as a proxy for near-Haar-uniform coverage of $U(d^k)$, the regime $0<\varepsilon\le 1/10$, and the requirement $N_{nt}\ge 2/\|O\|$ [2311.01330].

An illustration is given for H$_2$ VQE in the STO-3G basis with $d=2$, $k=2$, $\|H\|\approx 1.1686$, and $\varepsilon$ chosen $\simeq 0.01$. For each circuit depth $n$, the paper computes $N_{nt}(n)$ and writes
$$
\log C(\varepsilon;n)\approx 16\,N_{nt}(n)\cdot \log(aN_{nt}(n)),
$$
with $a\approx 43$--$115$ depending on the upper or lower bound. Plotting average energy error $\Delta E$ against average expressibility $\log C$ yields a characteristic U-shaped curve. The low-expressibility side cannot reach the true ground state; the high-expressibility side suffers trainability issues, including barren plateaus; between them lies a plateau of depths called the set of acceptable points, or "best expressive region." The width of this region in expressibility units, $\Delta \log C$, is reported empirically to satisfy
$$
\Delta \log C\propto 1/\langle\Delta E\rangle
$$
[2311.01330].

The relation between [2606.12211] and [2311.01330] is therefore conceptual rather than identical. The former develops an information-theoretic Occam theory for learning unknown quantum states from copies, with adaptive oracle inequalities and matching sample lower bounds. The latter studies ansatz expressibility through two-sided covering-number bounds for observable-output hypothesis spaces and uses those bounds to identify an intermediate regime for ansatz design. This suggests that "sample-supported expressibility law" names a broader family of covering-based constraints on usable expressive capacity, but the precise operational meaning depends on whether the task is state learning from copies or variational optimization over parameterized circuits.

Source: https://www.emergentmind.com/topics/sample-supported-expressibility-law