---
title: Variance Allocation Problem
url: https://www.emergentmind.com/topics/variance-allocation-problem
type: topic
---

# Variance Allocation Problem

Searching arXiv for the cited papers to ground the article in current arXiv records.
The variance allocation problem denotes a family of optimization, attribution, and control problems in which variance, covariance, sampling dispersion, or uncertainty is distributed, decomposed, or regulated across components under a global objective or constraint. In portfolio theory, it appears as the allocation of the variance of a total return process to individual assets through a cooperative game [1606.09424]. In principal component analysis, it appears as the correct allocation, or misallocation, of between-condition variance across rotated components [1804.07079]. In survey sampling and statistical design, it appears as the choice of stratum or coordinate sample sizes to satisfy variance or coefficient-of-variation targets with minimum cost or minimum total sample size [1309.6148] [2204.04035] [1509.09286]. In quantum algorithms, it appears as shot redistribution across commuting measurement groups to reduce estimator variance [2507.16879]. This suggests a common mathematical motif: a variance-related resource or burden is assigned to units such as assets, components, strata, arms, tasks, or measurement groups so that fairness, efficiency, robustness, or inferential precision is attained.

## 1. Scope and recurring mathematical forms

Across the literature, the phrase refers to several distinct but structurally related problem classes.

| Domain | Decision object | Core criterion |
|---|---|---|
| Portfolio variance [1606.09424] | asset-level allocation \(\phi_i(\nu)\) | allocate \([S_N]\) via covariance with the total portfolio |
| PCA with conditions [1804.07079] | allocation of between-condition variance to components | keep the condition effect on one component |
| Stratified sampling [1309.6148] [2204.04035] | stratum sample sizes | meet CV or variance targets with minimum sample size or cost |
| ADAPT-VQE [2507.16879] | shots per commuting clique | minimize estimator variance or total shots |
| Gaussian variance budget [2502.18463] | \(\sigma_i^2\) or \(\Sigma\) | maximize an expected maximum |

Two recurring formulations are especially prominent. The first is a **fixed-budget allocation** problem: examples include \(\sum_{i=1}^n \sigma_i^2=1\) in Gaussian expectation maximization [2502.18463], \(\sum_{i=1}^m N_i=N\) in variance-based shot allocation [2507.16879], and \(\sum_{h=1}^H n_h=n\) or \(\sum_h c_h x_h\) in sampling design [1309.6148] [2204.04035]. The second is an **attribution or decomposition** problem, in which an already realized aggregate variance is assigned to constituents, as in the Shapley allocation of portfolio variance [1606.09424] and the predictable-process decomposition of terminal variance in continuous-time risk budgeting [2011.10747].

A related but distinct use studies the variability of the allocation itself. In stochastic multi-armed bandits, allocation variability is defined as
\[
S_T \triangleq \max_{i\in[K]} \sqrt{\mathrm{Var}(N_{i,T})},
\]
the largest standard deviation of any arm’s pull count [2602.07472]. This is not an allocation of variance to units; it is a variance measure of the allocation rule.

## 2. Portfolio-theoretic allocation and risk attribution

The canonical variance allocation problem in finance is formulated in the Markowitz mean-variance setting. If the portfolio return is
\[
S_N=\sum_{i=1}^{n}X_i,
\]
then correlated asset returns make the standalone quantity \([X_i]-\theta[X_i]\) inadequate as a contribution measure, because covariance creates interaction effects [1606.09424]. The construction therefore passes through a cooperative game with characteristic function
\[
\nu(J)=\left[\sum_{i\in J}X_i\right].
\]
The Shapley value is selected because it satisfies efficiency, symmetry, dummy player, and linearity, and because it interprets each asset’s allocation as its average marginal contribution over coalitions [1606.09424].

The central theorem gives a closed form:
\[
\phi_i(\nu)=[X_i,S_N]=\sum_{j=1}^{n}[X_i,X_j].
\]
Thus each asset is assigned exactly its covariance with the total portfolio return [1606.09424]. The utility allocation then becomes
\[
\phi_i(\gamma)=[X_i]-\theta[X_i,S_N].
\]
This rule has several notable implications. Assets with large positive covariance with the portfolio bear more variance cost; hedging assets can receive negative variance allocations; if total portfolio variance is zero under perfect hedging, then the Shapley vector is zero; and if all pairwise covariances are nonnegative, the variance game is supermodular, whereas if they are nonpositive, it is submodular [1606.09424]. Computationally, the result replaces a generic factorial-time Shapley computation by the row sums of the covariance matrix, yielding a quadratic-time calculation [1606.09424].

Later continuous-time work generalizes the same intuition from static covariance matrices to stochastic processes. Using terminal variance as the risk measure, the marginal risk contribution is represented by a predictable density process \(c(t,\omega)\), and the asset-wise risk contribution becomes
\[
u^{(i)}(t,\omega)c^{(i)}(t,\omega),
\]
so that total risk is aggregated over \([0,T]\times\Omega\) [2011.10747]. The associated risk budgeting problem seeks \(u^\star\) such that
\[
u^\star_t\odot c^{u^\star}_t=\beta_t,
\]
turning static Euler risk attribution into a stochastic process identity [2011.10747].

Other mean-variance formulations alter what is being allocated. Tracking-error penalization adds the running term
\[
\mathbb{E}\!\left[\int_0^T (\alpha_t-w_r X_t)^\top \Gamma (\alpha_t-w_r X_t)\,dt\right]
\]
to the mean-variance criterion, and when \(\Gamma=\gamma I_d\) with \(\gamma\to\infty\), the optimal policy converges to the reference portfolio \(w_r X_t\) [2009.08214]. In general incomplete markets, equilibrium mean-variance policy decomposes into a myopic term and a hedging term [2412.18498]. In exploratory reinforcement learning, the optimal policy is multivariate Gaussian with covariance
\[
(\sigma'\sigma)^{-1}\frac{\lambda}{2}e^{\rho'\rho(T-t)},
\]
so policy variance itself becomes a control variable rather than merely an outcome [1907.11718].

## 3. Component models, PCA, and variance misallocation

In PCA with treatment or condition factors, the variance allocation problem concerns whether between-condition variance is placed on the correct component or spread across several components after rotation [1804.07079]. Following Wood and McCarthy (1984), the ideal is that a single PCA component carries the complete between-condition variance of a condition factor together with some within-condition variance [1804.07079].

The paper establishes that rotation is the mechanism of misallocation. If, before rotation, only the first component has a nonzero condition mean,
\[
E(w_{1i})\neq 0,\qquad E(w_{ji})=0 \text{ for } j>1,
\]
then after a nontrivial rotation the condition effect generally appears on components other than the first as well [1804.07079]. Hence a solution that initially localizes the effect on one component can be made ambiguous by rotation alone.

The decisive structural issue is not loading magnitude but loading shape. Theorems 2–4 show progressively weaker conditions for unambiguous allocation. If the within-condition loading matrices satisfy
\[
A_i^v=A^b \quad \text{for all } i,
\]
then within- and between-condition parts can be combined without misallocation [1804.07079]. For the single between-condition factor case \(q^b=1\), it is enough that one within-condition loading vector matches the between-condition loading vector,
\[
a_1^v=a^b,
\]
so that the first component combines within- and between-condition variance cleanly [1804.07079]. Most importantly, exact equality of magnitudes is not required. If the loading vectors are proportional,
\[
\theta_s a_s^v = a^b,\qquad \theta_s>0,
\]
then the condition variance is still allocated optimally to one component [1804.07079].

This yields a specific correction to a common misconception. Different loading magnitudes across conditions do not by themselves imply variance misallocation. The necessary condition is similar loading shape, not similar loading magnitude [1804.07079]. For the same reason, perfect Tucker congruence is not required, whereas a perfect Pearson correlation between corresponding loading vectors is sufficient because it captures identical shape up to linear scaling [1804.07079].

## 4. Sample-size, survey, and measurement-budget allocation

In survey methodology, variance allocation is usually the problem of choosing stratum sample sizes so that estimator precision is achieved with minimum sample size or minimum cost. For multivariate stratified sampling with study variables \(Y_1,\dots,Y_m\), the variance of the estimator of the total for variable \(Y_j\) is
\[
V(\hat t_{y_j}) = \sum_{h=1}^H \frac{N_h^2 S_{hj}^2}{n_h}\left(1-\frac{n_h}{N_h}\right),
\]
so the allocation vector \((n_h)\) directly controls precision [1309.6148]. One exact approach introduces binary variables
\[
x_{hk}=1 \iff \text{sample size } k \text{ is chosen for stratum } h,
\]
and solves a pure binary integer program that minimizes total sample size while enforcing coefficient-of-variation constraints for every study variable [1309.6148]. This avoids the noninteger solutions and post hoc rounding issues of nonlinear approximations [1309.6148].

A complementary line of work studies optimum allocation in stratified sampling through Karush-Kuhn-Tucker conditions. With generic variance function
\[
V_{\hat\theta}(x)=\sum_{h\in\mathcal H}\frac{A_h^2}{x_h}-A_0,
\]
the problem of minimizing cost subject to a fixed variance level and upper bounds on stratum sizes is transformed into a convex lower-bounded problem in variables \(z_h=A_h^2/(c_hx_h)\) [2204.04035]. The resulting algorithm, Recursive Neyman Allocation under lower bounds (LRNA), is proved optimal and positioned as the lower-bound counterpart of recursive Neyman allocation with upper bounds [2204.04035].

When the stratum variances are themselves estimated rather than known, the objective vector becomes random. The multivariate stratified allocation problem is then formulated as an integer nonlinear stochastic multiobjective program, with asymptotic normal approximation for the vector of sample variances and solution strategies based on E-model, V-model, P-model, and Kataoka formulations [1106.0773]. This replaces a single deterministic variance criterion by an explicitly stochastic multiobjective one [1106.0773].

Nonresponse introduces another design-stage variance allocation problem. Expected-response-rate allocation sets
\[
n_h^{ERR}= \left(\frac{1}{r_h}\frac{N_h}{\sum_{i=1}^{H}N_i}m\right),
\]
whereas proportional-to-size allocation uses the overall average response rate \(r\) instead [2005.12168]. The resulting asymptotic variance under ERR is
\[
\sigma^{2^{ERR}(\hat q) = \frac{1}{Nm}\sum_{h=1}^{H}N_h q_h(1-q_h)\frac{r_h}{p_h},
\]
and when the expected response rates are correctly specified, \(r_h=p_h\), ERR yields no larger variance than proportional-to-size allocation [2005.12168].

A related statistical formulation appears in the many-normal-means model with heterogeneous coordinate noise. There the decision variables are the coordinate measurement counts \(n_i\) under a total budget constraint \(\sum_i n_i\le n\), and the objective is the minimax linear risk over ellipsoids or hyperrectangles [1509.09286]. For ellipsoids, an explicit suboptimal allocation has the form
\[
n_{s,i}\propto \sigma_i[1-ta_i^{-1}]_+^{1/2},
\]
while for hyperrectangles the exact optimal allocation has a finite active set and can satisfy \(n_{o,i}=0\) for \(i>d_o\) [1509.09286]. In both cases, reallocation improves on uniform allocation, and for ellipsoids the paper explicitly states that it improves the Pinsker bound [1509.09286].

## 5. Sequential experimentation, quantum measurements, and adaptive allocation

In sequential experimentation, the variance allocation problem shifts from decomposing an existing variance to distributing future samples so that estimator variance or decision error is reduced. In ADAPT-VQE, the costly measurements are Hamiltonian expectation estimation and operator-gradient estimation. The observable is partitioned into commuting cliques, and variance-minimized shot assignment (VMSA) solves the fixed-budget problem by allocating the remaining shots proportionally to estimated clique standard deviations after a pilot sample [2507.16879]. Variance-preserved shot reduction (VPSR), following Zhu et al., instead minimizes total shots subject to a target variance threshold [2507.16879]. The same paper combines variance-based shot allocation with reuse of Pauli measurements from the previous VQE step, reporting shot reductions required to reach chemical accuracy of \(6.71\%\) and \(43.21\%\) on H\(_2\) for VMSA and VPSR, and \(5.77\%\) and \(51.23\%\) on LiH, while maintaining chemically accurate energies where achievable [2507.16879].

Fixed-budget best-arm identification with heterogeneous variances uses the same principle at the arm level. With known reward variances, SHVar pulls the arm maximizing
\[
\frac{\sigma_i^2}{N_{s,t,i}},
\]
which approximately yields stage allocations
\[
N_{s,i}\approx \frac{\sigma_i^2}{\sum_{j\in A_s}\sigma_j^2}\,n_s
\]
and equalizes the variances of the empirical means [2306.07549]. The paper explicitly relates this to G-optimal design [2306.07549]. With unknown variances, SHAdaVar replaces \(\sigma_i^2\) by an upper confidence bound \(U_{s,t,i}\) and samples the arm with the largest optimistic estimate of sample-mean variance [2306.07549].

Unknown sampling variance also changes classical simulation-budget allocation in ranking and selection. In the Bayesian formulation with unknown means and variances, the exponential decay rate of \(1-PCS_B^n\) is
\[
\min_{i\neq i^*} V_i(\alpha_i,\alpha_{i^*}),
\]
and the pairwise objective is nonconvex because the auxiliary minimizer can be discontinuous as the allocation ratio changes [2509.02138]. This distinguishes the unknown-variance case from known-variance OCBA-type equations. The sequential procedure \(\mathcal{OCBA}^{\mathcal U}\) is then shown to learn the optimal allocation asymptotically without tuning parameters or forced exploration [2509.02138].

A related but conceptually different result concerns the variance of the allocation rule itself. In stochastic bandits, any active-learning algorithm with sublinear worst-case regret must satisfy
\[
\mathcal R_T\cdot \mathcal S_T=\Omega(T^{3/2}),
\]
where \(\mathcal S_T\) is worst-case allocation variability [2602.07472]. Hence minimax regret-optimal algorithms necessarily have worst-case allocation variability \(\Theta(T)\), the largest possible scale [2602.07472]. This shows that efficient learning and stable allocation are incompatible in that model.

## 6. Variance as a resource, a control target, and a fairness surrogate

Some recent work treats variance as a resource that can be deliberately placed where it yields the highest return. In Gaussian expectation maximization, the decision variables are the marginal variances or, in the correlated version, the covariance matrix, under the budget constraint
\[
\sum_{i=1}^n \sigma_i^2=1
\quad\text{or}\quad
\sum_{i=1}^n \Sigma_{ii}=1,\ \Sigma\succeq 0.
\]
The objective is to maximize \(\mathbb E[\max_i X_i]\) or \(\mathbb E[\sum_j \max_{i\in S_j} X_i]\) [2502.18463]. The structural conclusion is that optimal variance allocation concentrates on a small subset of variables as \(|S_j|\) increases, with a PTAS for the single-set case and an \(O(\log n)\)-approximation for the general multi-set case [2502.18463]. Concentration is not universal, however: for the cycle instance \(S_j=\{X_j,X_{j+1}\}\) with equal means, the optimal solution is uniform,
\[
\sigma_i^2=\frac{1}{n}\quad \forall i
\]
[2502.18463].

In multi-robot task allocation, the ensemble state is modeled as a stochastic jump process, and the transition-rate structure is chosen so that the mean depends only on \(k_{ij}\), whereas the second moment depends on \(\beta_i\) as well [2212.09816]. This decouples mean regulation from covariance shaping. The paper states that larger \(\beta_i\) implies smaller steady-state covariance, assuming rates remain positive, and illustrates this by reducing variances in a four-task example from
\[
[5.78,\ 6.83,\ 4.20,\ 1.44]'
\]
to
\[
[1.06,\ 1.12,\ 1.15,\ 0.45]'
\]
while keeping the mean close to the target [2212.09816].

In grouped-subcarrier OFDMA, variance is used as a scheduling metric rather than a risk measure. Users whose reported group gains have larger variance are given priority in the first allocation stage, because they benefit more from early access to their best groups [1207.4973]. Conflicts are resolved by descending-order variance assignment, and remaining groups are allocated through a fairness-enhancement stage using the criterion
\[
\min \frac{R_k}{\alpha_k}.
\]
The reported modified Jain fairness index is around \(0.99\) [1207.4973].

Fair division yields a cautionary counterexample. Minimizing the sum of variances of realized values subject to ex-ante proportionality can work well when valuations are identical: every allocation in the support is EFX, implying \(4/7\)-MMS for \(n\ge 4\) and \(2/3\)-MMS for \(n=3\) [2601.16579]. But when valuations are not identical, the same approach can fail even for two agents and two goods: the variance-minimizing ex-ante proportional distribution may assign both goods to one agent with positive probability, so the support need not even be EF1 [2601.16579]. This sharply limits the use of variance minimization as a proxy for ex-post fairness.

Across these literatures, variance is alternately a cost, a budget, a decomposition target, a control state, a sampling proxy, and a fairness surrogate. The shared theme is not a single canonical optimization model but a recurrent structural problem: how to place, apportion, or regulate dispersion so that a system-level objective is optimized without ignoring interaction effects.

Source: https://www.emergentmind.com/topics/variance-allocation-problem