---
title: Average Sample Complexity
url: https://www.emergentmind.com/topics/average-sample-complexity
type: topic
---

# Average Sample Complexity

Average sample complexity is not a single universally standardized object; in the literature represented here, it denotes several technically distinct notions that share an averaging principle over observations, samples, or sampled time sets. In reinforcement learning and stochastic optimization, it refers to the number of samples needed to produce an \(\varepsilon\)-accurate policy, value estimate, or optimizer, typically under a generative model or Sample Average Approximation (SAA) framework. In dynamical systems, it is a cover- or partition-level invariant that averages complexity over subsets of a finite orbit window rather than over the full window alone [1512.01143; 2509.20738]. A narrower algorithmic usage also appears in generalized Leapfrogging Samplesort, where a sorted sample is used recursively to partition an unsorted block, yielding an average-case recurrence
\[
A(n)=A(s)+(2k-1)(s+1)\log(s+1)+(s+1)A(2k-1),
\]
with \(n=s+(2k-1)(s+1)\), and hence \(A(n)=O(n\log n)\) for constant \(k\) [1801.09431]. This suggests that the term is best understood through its local context: statistical learning, stochastic optimization, dynamical complexity, or sample-based partitioning.

## 1. Terminological scope and mathematical forms

In average-reward Markov decision processes (MDPs), average sample complexity is tied to the task of learning an \(\varepsilon\)-optimal policy from samples drawn by a generative model. A representative objective is to output \(\hat\pi\) such that
\[
0\le \bar\alpha-\alpha^{\hat\pi}\le \epsilon
\]
with probability at least \(1-\delta\), while minimizing the total number of queried transitions [2310.08833]. In constrained and robust variants, the target is modified to include feasibility conditions or worst-case transitions, but the operative quantity remains the sample budget required for a prescribed error tolerance [2509.16586; 2505.10007].

In stochastic programming, the analogous role is played by SAA. One replaces an expectation \(F(x)=\mathbb E[f(x,\xi)]\) by an empirical objective \(F_N(x)\), or by a nested empirical estimator in conditional stochastic optimization, and then asks how large \(N\) must be so that an empirical minimizer is \(\varepsilon\)-optimal in expectation or with high probability [2401.00664; 1905.11957]. In this setting, “average” refers to empirical averaging over random samples rather than long-run average reward.

In dynamical systems, by contrast, average sample complexity is itself the invariant. For a topological system \((X,T)\) and open cover \(\mathscr U\), it is defined by
\[
\Asc(X,\mathscr U,T):=\lim_{n\to\infty}\frac{1}{n}\sum_{S\subset n^*} c_S^n \log N(\mathscr U_S),
\]
while for a measure-preserving system and finite partition \(\alpha\),
\[
\Asc_\mu(X,\alpha,T):=\lim_{n\to\infty}\frac{1}{n}\sum_{S\subset n^*} c_S^n H_\mu(\alpha_S).
\]
Its companion quantity, intricacy, satisfies
\[
\Int=2\,\Asc-h
\]
in both topological and measure-theoretic forms [1512.01143].

## 2. Average-reward MDPs and the emergence of optimal rates

A central modern use of average sample complexity is in average-reward reinforcement learning under a generative model. Early work established a reduction from average-reward MDPs to discounted MDPs under a mixing-time assumption: for every deterministic stationary policy, the induced chain has mixing time at most \(t_{\mathrm{mix}}\). Under this regime, an \(\epsilon\)-optimal deterministic stationary policy can be found using
\[
\widetilde O\!\left(\frac{A_{\mathrm{tot}}\,t_{\mathrm{mix}}^3}{\epsilon^3}\right)
\]
oblivious samples per state-action pair, while any algorithm using deterministic oblivious samples requires at least
\[
\Omega\!\left(\frac{A_{\mathrm{tot}}\,t_{\mathrm{mix}}^2}{\epsilon^2}\right)
\]
in the worst case [2106.07046].

Subsequent work closed the principal gap for uniformly ergodic average-reward MDPs. Using the minorization parameter
\[
\mathrm{minorize}:=\max_{\pi\in\Pi}\mathrm{minorize}(P_\pi),
\]
with \(\mathrm{minorize}\asymp t_{\mathrm{mix}}\) up to universal constants, the optimal policy-learning complexity became
\[
\widetilde\Theta\!\left(\frac{|S||A|\,\mathrm{minorize}}{\epsilon^2}\right)
\]
or equivalently
\[
\widetilde\Theta\!\left(\frac{|S||A|\,t_{\mathrm{mix}}}{\epsilon^2}\right),
\]
thereby matching the literature’s lower bound up to logarithmic factors [2310.08833].

A different line of work replaced uniform mixing by the span of the optimal bias function,
\[
H:=\|h^\star\|_{\mathrm{span}},
\]
and proved, for weakly communicating MDPs,
\[
\widetilde O\!\left(SA\frac{H}{\varepsilon^2}\right)
\]
samples under a generative model. This is minimax optimal up to logarithmic factors in \(S\), \(A\), \(H\), and \(\varepsilon\), and it is strictly more general than guarantees based on a uniform mixing-time bound because \(H\) can remain finite when uniform mixing is not available [2311.13469]. The multichain extension showed that span alone is insufficient in general average-reward MDPs; a bounded transient-time parameter
\[
\mathbb E_s^\pi[T_{R^\pi}]\le B \qquad \forall \pi,s
\]
is necessary, leading to the optimal bound
\[
\widetilde O\!\left(SA\frac{B+H}{\varepsilon^2}\right)
\]
with a matching minimax lower bound up to logarithmic factors [2403.11477].

The plug-in approach, long regarded as the simplest model-based baseline, was subsequently given its first finite-sample analysis. With empirical transitions, optional anchoring, and exact average-reward planning in the estimated model, it achieves the optimal diameter-based and mixing-based rates
\[
\widetilde O\!\left(SA\frac{D}{\varepsilon^2}\right),\qquad
\widetilde O\!\left(SA\frac{\tau_{\mathrm{unif}}}{\varepsilon^2}\right)
\]
without prior knowledge of \(D\) or \(\tau_{\mathrm{unif}}\), and also admits span-based bounds with unavoidable empirical-span dependence [2410.07616].

The main average-reward sample-complexity laws established in these works can be organized as follows.

| Setting | Complexity law | Structural parameter |
|---|---:|---|
| Uniformly ergodic AMDP | \(\widetilde\Theta(|S||A|t_{\mathrm{mix}}\epsilon^{-2})\) | \(t_{\mathrm{mix}}\) / minorize |
| Weakly communicating AMDP | \(\widetilde O(SA H \varepsilon^{-2})\) | \(H=\|h^\star\|_{\mathrm{span}}\) |
| General multichain AMDP | \(\widetilde O(SA(B+H)\varepsilon^{-2})\) | \(B+H\) |
| Plug-in, communicating or uniformly mixing | \(\widetilde O(SA D \varepsilon^{-2})\), \(\widetilde O(SA\tau_{\mathrm{unif}}\varepsilon^{-2})\) | \(D\), \(\tau_{\mathrm{unif}}\) |

These results collectively show that the “right” average sample complexity parameter depends on the structural class: uniform ergodicity favors \(t_{\mathrm{mix}}\), weak communication favors \(H\), and multichain structure requires \(B\) in addition to \(H\) [2310.08833; 2311.13469; 2403.11477; 2410.07616].

## 3. Constrained, robust, offline, and model-free extensions

Constrained average-reward MDPs introduce long-run feasibility conditions. For a stationary policy \(\pi\),
\[
\rho_r^\pi(s)=\lim_{T\to\infty}\frac{1}{T}\mathbb E_s^\pi\!\left[\sum_{t=0}^{T-1} r(s_t,a_t)\right],\qquad
\rho_c^\pi(s)=\lim_{T\to\infty}\frac{1}{T}\mathbb E_s^\pi\!\left[\sum_{t=0}^{T-1} c(s_t,a_t)\right],
\]
and the goal is
\[
\max_\pi \ \rho_r^\pi \quad \text{s.t.}\quad \rho_c^\pi\ge b.
\]
Under a generative model, model-based primal-dual learning yields a sharp separation between relaxed and strict feasibility:
\[
\tilde O\!\left(\frac{SA(B+H)}{\epsilon^2}\right)
\]
samples for relaxed feasibility and
\[
\tilde O\!\left(\frac{SA(B+H)}{\epsilon^2\zeta^2}\right)
\]
for strict feasibility, where
\[
\zeta:=\max_\pi \rho_c^\pi-b
\]
is the Slater constant. A matching lower bound
\[
\tilde\Omega\!\left(\frac{SA(B+H)}{\epsilon^2\zeta^2}\right)
\]
shows that the \(\zeta^{-2}\) penalty is unavoidable when exact feasibility is required [2509.16586].

Robust and distributionally robust variants alter the transition model rather than adding explicit constraints. For robust average-reward policy evaluation under contamination, total variation, and Wasserstein uncertainty sets, the robust Bellman operator is a contraction under the span seminorm,
\[
\|v\|_{\mathrm{sp}}=\max_s v(s)-\min_s v(s),
\]
and truncated Multi-Level Monte Carlo (MLMC) yields the first finite-sample guarantee:
\[
\tilde{\mathcal O}(\epsilon^{-2})
\]
for both robust value estimation and robust average reward estimation, with explicit dependence on \(S\), \(A\), \(t_{\mathrm{mix}}\), \((1-\gamma)^{-1}\), and the uncertainty radius \(\delta\) [2502.16816]. For distributionally robust average-reward reinforcement learning under KL and \(f_k\)-divergence uncertainty sets, two algorithms—one based on reduction to a robust discounted MDP and one based on anchoring—both achieve
\[
\widetilde{O}\!\left(|S||A|^2 t_{\mathrm{mix}}^2\varepsilon^{-2}\right)
\]
under uniform ergodicity and sufficiently small uncertainty radius, providing the first finite-sample convergence guarantee for this setting [2505.10007].

Offline average-reward RL introduces nonuniform coverage effects. In weakly communicating MDPs, a single-policy guarantee is possible in terms of the target policy’s bias span and policy hitting radius. If \(\pi\) is deterministic and unichain and the tabular offline data satisfy
\[
n(s,\pi(s)) \ge m\mu^\pi(s)+\alpha\big(C_2(P,\pi)\big)^2+4 \qquad \forall s\in\mathcal S,
\]
then with high probability the returned policy obeys
\[
\rho^\pi \ge \rho^\star - \sqrt{\frac{C_1\,S\,(\|h^\pi\|_{\mathrm{span}}+1)\,\alpha}{m}}.
\]
The additive transient-coverage term is essential; stationary-distribution coverage alone is insufficient [2506.20904].

Model-free average-reward Q-learning remains statistically less efficient than the best model-based rates, but recent work gives explicit guarantees. Under weakly communicating MDPs, single-agent Q-learning with stage-wise discount scheduling achieves
\[
\widetilde O\!\left(\frac{|\mathcal S||\mathcal A|\|h^\star\|_{\mathsf{sp}}^3}{\varepsilon^3}\right),
\]
while a federated version with \(M\) agents reduces the per-agent sample complexity to
\[
\widetilde O\!\left(\frac{|\mathcal S||\mathcal A|\|h^\star\|_{\mathsf{sp}}^3}{M\varepsilon^3}\right)
\]
and requires only
\[
\widetilde O\!\left(\frac{\|h^\star\|_{\mathsf{sp}}}{\varepsilon}\right)
\]
communication rounds [2601.13642].

## 4. Sample Average Approximation and stochastic optimization

In stochastic programming, average sample complexity is usually expressed through the sample size required by SAA. For conditional stochastic optimization,
\[
\min_{x\in\mathcal X} F(x)
\quad\text{where}\quad
F(x):=\mathbb E_{\xi}\!\left[f_\xi\!\left(\mathbb E_{\eta\mid \xi}[g_\eta(x,\xi)]\right)\right],
\]
the nested empirical objective
\[
\hat F_{nm}(x)=\frac1n \sum_{i=1}^n f_{\xi_i}\!\left(\frac1m\sum_{j=1}^m g_{\eta_{ij}}(x,\xi_i)\right)
\]
is generally biased. Under Lipschitz assumptions, the total sample complexity is
\[
\mathcal O(d/\varepsilon^4),
\]
it improves to
\[
\mathcal O(d/\varepsilon^3)
\]
when the outer function is smooth, and under a Hölderian error bound it further improves to
\[
\mathcal O(1/\varepsilon^2)
\]
in the smooth quadratic-growth case. In the special case where \(\xi\) and \(\eta\) are independent and one uses a modified SAA with shared inner samples, the total complexity becomes
\[
\mathcal O(d/\varepsilon^2)
\]
[1905.11957].

For convex and strongly convex stochastic programming, metric entropy-free analyses have materially changed the status of SAA. In the strongly convex case, if
\[
N\ge C_1 \max\left\{\frac{(\sigma_p^2+M^2)}{\varepsilon^2},\ \frac{C+M}{\varepsilon}\right\},
\]
then any SAA solution satisfies
\[
\mathbb E[F(x)-F(x^\star)]\le \varepsilon,
\]
and the high-probability counterpart introduces only a \(\ln(1/\beta)\) factor. These rates contain no covering-number or metric-entropy term and therefore match the canonical stochastic mirror descent rate up to minor logarithmic factors [2401.00664].

When sampling is biased because the random input is itself approximated, the complexity depends on the weak-error order. In biased SAA with
\[
\mathbb E[f(x,\zeta_h)] = \mathbb E[f(x,\zeta)] + c_1 h^\alpha + o(h^\alpha),
\]
standard Monte Carlo yields total cost
\[
\mathcal C=\mathcal O\!\left((d+\gamma)\log(\epsilon^{-1})\,\epsilon^{-(2+1/\alpha)}\right).
\]
By introducing MLMC into the SAA objective, one obtains
\[
\mathcal{C}_{\mathrm{mlmc}}^{\mathrm{saa}}=
\begin{cases}
\mathcal O\!\left((\gamma+d)\epsilon^{-2}\log(\epsilon^{-1})\right), & \beta>1,\\[4pt]
\mathcal O\!\left((\gamma+d)\epsilon^{-\left(2+\frac{\log(2)}{\alpha\log(m)}\right)}\log(\epsilon^{-1})\right), & \beta=1,\\[4pt]
\mathcal O\!\left((\gamma+d)\epsilon^{-\left(2+\frac{1-\beta}{\alpha}\right)}\log(\epsilon^{-1})\right), & \beta<1,
\end{cases}
\]
so MLMC recovers the unbiased Monte Carlo rate when \(\beta>1\) [2407.18504].

Distributionally robust SAA over a \(\phi\)-divergence ball exhibits a sharp dichotomy. If
\[
\lim_{x\to\infty}\frac{\phi(x)}{x}=+\infty,
\]
then a \(P\)-independent sample complexity exists and can be bounded through a growth-controlled factor \(M_{\phi,\tau}(\epsilon)\); for example, if
\[
n \ge \frac{128\,M_{\phi,\tau}(\epsilon)^2 \log\!\Big( \frac{32B^2\,\overline{L}_{\phi,\tau}(\epsilon)}{\tau\epsilon\delta} \Big)}{\epsilon^2},
\]
then the truncated SAA estimator attains \(\epsilon\)-accuracy with probability at least \(1-\delta\). If superlinear growth fails, then every estimator has a \(P\)-dependent lower bound that can be made arbitrarily large by changing the nominal law \(P\) [2604.10855].

## 5. Dynamical-systems average sample complexity

In topological and measure-preserving dynamics, average sample complexity is not a sample-size bound but a dynamical invariant. For a coefficient system \(\{c_S^n\}\) satisfying
\[
c_S^n\ge 0,\qquad \sum_{S\subset n^*}c_S^n=1,\qquad c_S^n=c_{S^c}^n,
\]
the topological quantity is
\[
\Asc(X,\mathscr U,T)
=\lim_{n\to\infty}\frac{1}{n}\sum_{S\subset n^*} c_S^n \log N(\mathscr U_S),
\]
and the measure-theoretic counterpart is
\[
\Asc_\mu(X,\alpha,T)
=\lim_{n\to\infty}\frac{1}{n}\sum_{S\subset n^*} c_S^n H_\mu(\alpha_S).
\]
The limits exist by subadditivity, and in both settings the associated intricacy satisfies
\[
\Int=2\,\Asc-h.
\]
Moreover,
\[
\sup_{\mathscr U}\Asc(X,\mathscr U,T)=h_{\mathrm{top}}(X,T),\qquad
\sup_{\alpha}\Asc_\mu(X,\alpha,T)=h_\mu(X,T),
\]
so the supremum over all covers or partitions recovers ordinary entropy [1512.01143].

These quantities nonetheless discriminate between systems at fixed cover or partition. For a full \(r\)-shift,
\[
\Asc(\Sigma_r,\mathscr U_0,\sigma)=\frac12\log r,\qquad
\Int(\Sigma_r,\mathscr U_0,\sigma)=0,
\]
while for certain subshifts of finite type with adjacency matrix \(M\) satisfying \(M^2>0\),
\[
\Asc(X,\mathscr U_0,\sigma)=\frac14\sum_{k=1}^\infty \frac{\log |\mathscr L_{k^*}(X)|}{2^k}.
\]
The measure-theoretic theory also yields explicit formulas for Markov shifts, such as
\[
\Asc_\mu(X,\alpha,T)=\frac12\sum_{i=1}^\infty 2^{-i} H_\mu(\alpha\mid\alpha_i)
\]
for a \(1\)-step Markov shift with the time-\(0\) partition [1512.01143].

The amenable-group extension introduces local and relative versions. For a factor map \(\pi:(X,G)\to(Y,G)\), finite cover \(\mathcal U\), Følner sequence \(\{F_n\}\), and uniform coefficients \(c_S^{F_n}=2^{-|F_n|}\),
\[
\mathrm{Asc}_{\mathrm{top}}(G,\mathcal U\mid Y)
=
\lim_{n\to\infty}
\frac{1}{|F_n|}
\sum_{S\subseteq F_n}
c_S^{F_n}\log N(\mathcal U_S\mid Y),
\]
and similarly for the measure-theoretic relative quantity with \(H_\mu(\alpha_S\mid Y)\). The paper proves the equivalence
\[
\mathrm{Asc}_\mu^-(G,\mathcal U)=\mathrm{Asc}_\mu^+(G,\mathcal U)
\]
for general amenable groups, and establishes a local variational principle:
\[
\mathrm{Asc}_{\mathrm{top}}(G,\mathcal U)
=
\max_{\mu\in\mathcal M(X,G)}\mathrm{Asc}_\mu(G,\mathcal U)
=
\max_{\mu\in\mathcal M^e(X,G)}\mathrm{Asc}_\mu(G,\mathcal U).
\]
In uniquely ergodic \(\mathbb Z\)-systems, the topological and the two measure-theoretic local versions coincide for open covers [2509.20738].

## 6. Structural parameters, separations, and persistent misconceptions

A recurring theme is that average sample complexity is controlled by structure, not just by ambient dimension or state-action cardinality. In average-reward MDPs, the operative parameter may be the uniform mixing time \(t_{\mathrm{mix}}\), the minorization constant, the optimal bias span \(H\), the diameter \(D\), or the transient time \(B\), depending on the class of MDP being studied [2310.08833; 2311.13469; 2403.11477; 2410.07616]. A common misconception is that span alone always characterizes average-reward difficulty. The multichain lower-bound constructions show otherwise: there exist instances with
\[
\|h^\star\|_{\mathrm{span}}=0
\]
that still require many samples because of long transient behavior, which is precisely why \(B\) appears in the general complexity law [2403.11477].

A second misconception is that exact feasibility or robustness merely changes constants. In constrained average-reward MDPs, strict feasibility incurs the unavoidable factor \(\zeta^{-2}\) absent from relaxed feasibility [2509.16586]. In offline average-reward RL, stationary coverage alone is not enough; the transient-coverage term is necessary [2506.20904]. In plug-in average-reward planning, a purely \(\operatorname{span}(h^\star)\)-based guarantee is impossible for the plug-in template because a lower bound forces additional empirical-span dependence [2410.07616]. In distributionally robust SAA, distribution-free sample complexity is possible only when the divergence function grows superlinearly; otherwise the lower bound is inherently nominal-distribution-dependent [2604.10855].

The dynamical-systems literature resolves a different misunderstanding. Since
\[
\sup \Asc = h
\]
at the system level, one might conclude that average sample complexity is merely entropy in disguise. The explicit formulas for symbolic systems and the relative cover-level theory show that this is not the case: at fixed cover, partition, or factor, average sample complexity detects structure that ordinary entropy does not isolate [1512.01143; 2509.20738].

Taken together, these results indicate that “average sample complexity” is best treated as a family of parameter-sensitive asymptotic laws or invariants. In learning and optimization, it quantifies the sample budget needed to achieve a target accuracy under average-reward, nested-expectation, or robust criteria. In dynamics, it measures the average informational content of sampled orbit subsets. The shared mathematical intuition is averaging over partial observations; the specific object being averaged, and the structural parameter governing it, are determined by the problem class.

Source: https://www.emergentmind.com/topics/average-sample-complexity