---
title: Exact Quantile Balancing (EQB) in Data Science and Research
url: https://www.emergentmind.com/topics/exact-quantile-balancing-eqb
type: topic
---

# Exact Quantile Balancing (EQB) in Data Science and Research

Exact Quantile Balancing (EQB) denotes methods that impose, exploit, or compute exact quantile constraints in distributed computation, causal inference, survey weighting, predictive calibration, and machine-learning systems. The term is not used with a single universal definition. In inverse-propensity-weighting sensitivity analysis, EQB is a sharp population and finite-sample calibration procedure that balances a conditional outcome quantile to recover the narrowest interval allowed by a marginal sensitivity model [2102.04543]. In non-probability sampling, it denotes calibration weights that exactly reproduce selected population quantiles of auxiliary variables [2403.09726]. In distributed and mixture-of-experts systems, it refers to exact computation of global empirical quantiles, or to their use for balancing assignments, rather than to reweighting observations [2511.12025; 2609.28053]. Related quantile-based procedures in classification, fairness, quantile regression, and causal quantile treatment-effect estimation provide adjacent techniques but do not, without qualification, establish exact finite-sample EQB [1803.00067; 1907.08646; 2109.03757].

## 1. Terminology and conceptual scope

A quantile is an order-statistic functional of a distribution. For a distribution function $F$, its quantile function is

$$
Q(\tau)=\inf\{y:F(y)\geq \tau\},
$$

where $\tau\in(0,1)$. In finite samples, exact quantile computation concerns an order statistic or a value satisfying the lower- and upper-rank condition

$$
\operatorname{rank}_{<}(x)<k\leq \operatorname{rank}_{\leq}(x).
$$

Exact quantile balancing adds a balancing requirement: selected quantile-related functionals must coincide with specified population, group, treatment, or assignment targets. Depending on the setting, the balanced object may be a covariate distribution, an outcome quantile, a terminal Gaussian statistic, or an expert-routing margin.

The word “exact” has several distinct meanings:

- **Exact population identification**: in sensitivity analysis, a quantile constraint characterizes the sharp endpoint of a partially identified causal parameter.
- **Exact finite-sample calibration**: in survey weighting, weights solve equations that reproduce selected population quantiles or their indicator representations.
- **Exact empirical computation**: in distributed systems, the returned order statistic is exact at the machine representation used by the algorithm.
- **Exact assignment balance**: in mixture-of-experts routing, an empirical quantile produces the bias required for a balanced global assignment problem, subject to rank, divisibility, and tie conventions.

These meanings should not be conflated with approximate or asymptotic quantile balancing. A method may use quantile-derived regressors, prove $O_p(1/\sqrt n)$ fairness, or estimate a weighted outcome distribution without imposing exact finite-sample quantile equations. “Constrained Classification and Ranking via Quantiles” uses score quantiles to construct thresholds associated with predicted positive or negative rates, but the supplied material does not establish exact finite-sample rate constraints [1803.00067]. Similarly, “Fair quantile regression” establishes $\sqrt n$-fairness rather than literal finite-sample equality [1907.08646].

## 2. Exact quantile balancing in inverse-propensity sensitivity analysis

The most formal use of EQB is in sensitivity analysis for inverse propensity weighting under Tan’s marginal sensitivity model. Observed data are $(X_i,Y_i,Z_i)$, with binary treatment $Z_i$, potential outcomes $Y_i(0),Y_i(1)$, nominal propensity score

$$
e(x)=P(Z=1\mid X=x),
$$

and unobserved confounding represented through a true propensity score $e_0(x,u)$. The model restricts the odds ratio between the true and nominal propensity scores:

$$
\Lambda^{-1}
\leq
\frac{e_0(X,U)/[1-e_0(X,U)]}{e(X)/[1-e(X)]}
\leq \Lambda .
$$

The sensitivity parameter satisfies $\Lambda\geq1$. When $\Lambda=1$, the true and nominal odds coincide; larger values permit stronger unobserved selection.

Ordinary sensitivity analysis optimizes stabilized inverse-propensity estimates over pointwise propensity constraints. The resulting ZSB interval can be asymptotically too wide because pointwise odds-ratio restrictions do not ensure compatibility with the observed distribution of $X$. Any compatible propensity must satisfy, for every integrable function $h$,

$$
E\left[\frac{h(X)Z}{e_0(X,U)}\right]=E[h(X)].
$$

Thus, the admissible propensity must balance infinitely many functions of $X$. EQB reduces these infinitely many restrictions, for the purpose of bounding counterfactual means, to one conditional outcome-quantile constraint for each endpoint.

Let

$$
Q_t(x,z)=\inf\{q:P(Y\leq q\mid X=x,Z=z)\geq t\},
$$

and set

$$
\tau=\frac{\Lambda}{\Lambda+1}.
$$

The sharp upper bound for the treated mean balances the conditional $\tau$-quantile:

$$
\psi_{\mathrm T}^{+}
=
\max_{\bar E\in\mathcal E_\infty(\Lambda)}
E\left[\frac{YZ}{\bar E}\right]
$$

subject to

$$
E\left[\frac{Q_{\tau}(X,1)Z}{\bar E}\right]
=
E[Q_{\tau}(X,1)].
$$

The sharp lower bound balances the conditional $(1-\tau)$-quantile:

$$
\psi_{\mathrm T}^{-}
=
\min_{\bar E\in\mathcal E_\infty(\Lambda)}
E\left[\frac{YZ}{\bar E}\right]
$$

subject to

$$
E\left[\frac{Q_{1-\tau}(X,1)Z}{\bar E}\right]
=
E[Q_{1-\tau}(X,1)].
$$

The same construction applies to control means by replacing $Z$ with $1-Z$ and $\bar E$ with $1-\bar E$. The sharp ATE interval is

$$
[\psi_{\mathrm{ATE}}^{-},\psi_{\mathrm{ATE}}^{+}]
=
[\psi_{\mathrm T}^{-}-\psi_{\mathrm C}^{+},
\psi_{\mathrm T}^{+}-\psi_{\mathrm C}^{-}].
$$

The term “exact” refers here to sharpness: with consistently estimated conditional quantiles, the EQB endpoints converge to the narrowest interval compatible with the observed distribution and Tan’s model. If the conditional quantile model is misspecified, the estimated interval remains conservative in the stated direction: it can be too wide, but not asymptotically too narrow.

The worst-case inverse propensity has a threshold form. For the upper treated-mean bound, observations above $Q_\tau(X,1)$ receive the largest permitted inverse-propensity weight, while observations below the threshold receive the smallest permitted weight. The threshold is selected to satisfy the data-compatibility condition

$$
E\left[\frac{Z}{\bar E_+}\mid X\right]=1.
$$

This converts the infinite-dimensional balancing problem into a conditional linear program whose solution is obtained by thresholding the outcome at a conditional quantile. The associated finite-sample estimator uses $\hat e$ and estimated conditional quantiles $\hat Q_\tau$ and imposes normalized balance equations. Feasibility is guaranteed by using the IPW target on the right-hand side rather than an unweighted sample average.

The optimization has a weighted quantile-regression dual. With $g(X)=(1,\hat Q_\tau(X,1))$, the upper-bound problem is equivalent to a weighted quantile regression using the check loss

$$
\rho_\tau(u)=u\{\tau-\mathbb I(u<0)\}.
$$

The fitted quantile-regression residual determines which observations receive the upper or lower allowed inverse-propensity weight. This establishes a direct computational connection between EQB and weighted quantile regression.

## 3. Quantile calibration and weighting

In survey inference and non-probability sampling, quantile-balancing inverse probability weighting extends ordinary IPW by incorporating known or estimated auxiliary-variable quantiles. The target population is $U=\{1,\ldots,N\}$, with study variable $y_k$, auxiliary variables $x_k$, non-probability inclusion indicator $R_k$, and a probability reference sample $S_B$ with design weights $d_k^B$ [2403.09726].

Ordinary IPW estimates the unknown inclusion propensity $\pi_k^A=P(R_k=1\mid x_k)$ and uses weights $1/\widehat\pi_k^A$. Its estimating equations do not generally reproduce population totals, auxiliary distributions, or auxiliary quantiles. Calibration weighting instead adjusts design or inverse-propensity weights to satisfy target equations.

For an auxiliary variable $x_j$ with known population quantile $Q_{x_j,\alpha}$, exact quantile calibration imposes

$$
\widehat F_{x_j,\mathrm{cal}}(Q_{x_j,\alpha})=\alpha.
$$

The empirical distribution can be represented using indicator or interpolated-indicator variables. With a balance vector $a_k$ containing a normalization component and quantile indicators, quadratic calibration solves

$$
\min_v\sum_{k\in S_B}d_k^B
G\left(\frac{v_k^B}{d_k^B}\right)
$$

subject to

$$
\sum_{k\in S_B}v_k^B a_k=T_a,
$$

where $T_a$ contains the population total and the target quantile probabilities. If the equations converge and are feasible, the resulting weights reproduce the selected quantile constraints exactly.

The paper distinguishes two QBIPW constructions:

- **QBIPW-MLE**: quantile-derived variables enter an augmented propensity model. This generally improves distributional similarity but does not guarantee exact reproduction of quantiles.
- **QBIPW-GEE**: inverse-propensity parameters solve explicit calibration equations. The resulting weights reproduce auxiliary totals and selected quantile-balance totals exactly when the equations converge.

For selected quantile levels, such as quartiles $\{0.25,0.50,0.75\}$ or deciles $\{0.1,0.2,\ldots,0.9\}$, the quantile variables act as piecewise-constant basis functions. Exact balance is restricted to the selected quantile points; it does not reproduce the entire continuous distribution unless a sufficiently rich collection of constraints is used.

The construction requires conditional ignorability,

$$
R_k\perp y_k\mid x_k,
$$

positivity, and conditional independence of non-probability inclusion indicators given the auxiliary variables. Feasibility additionally requires sufficient sample size, non-collinearity, variation in the samples, and target moments lying within the relevant convex support. Too many quantile constraints can produce nonidentification, numerical failure, or extreme weights.

The simulations reported for QBIPW use a finite population of size $N=100{,}000$. Under nonlinear selection and outcome relationships, decile-based QBIPW-GEE substantially reduces bias relative to ordinary IPW. In the Polish job-vacancy application, QBIPW-GEE estimates of the share of vacancies directed toward Ukrainian workers are approximately between $21.53\%$ and $21.74\%$, with an example estimate of $21.53\%$ and standard error $0.60$.

## 4. Related uses in prediction, fairness, and causal quantile effects

Quantile-derived thresholds can control empirical prediction rates. If scores on a subset $A$ are $s_1,\ldots,s_m$ and $t$ is an empirical lower quantile at level $c$, then approximately a fraction $c$ of observations satisfy $s_i\leq t$, while approximately $1-c$ satisfy $s_i>t$. This creates a connection between quantile estimation and predicted positive or negative rate constraints.

The framework in “Constrained Classification and Ranking via Quantiles” models a threshold as a quantile of scores over a designated subset:

$$
\theta_A(w)=\widehat q(f(A);c).
$$

Its quantile-dependent loss uses the estimated threshold inside a loss function. Kernel, interval, and lower-bound quantile estimators are analyzed under minibatch sampling, with deviations of order

$$
O\left(\sqrt{\frac{1}{b}\log\frac{1}{\delta}}\right).
$$

The supplied material supports quantile-based surrogate learning and stochastic-gradient optimization, but not exact equality of predicted rates. Exact finite-sample rate balancing would require explicit subset constraints, exact order-statistic thresholds, tie handling, and full-data evaluation.

“Fair quantile regression” considers a predictor trained without observing a protected attribute $A$. For an estimated quantile $\widehat q_\tau(X)$, the group-specific effective quantile is

$$
\widehat\tau_a
=
P\{Y\leq \widehat q_\tau(X)\mid A=a\}.
$$

The post-processing procedure fits a residual quantile regression on $A$ and produces

$$
\widetilde q_\tau(a,x)
=
\widehat q_\tau(x)+\widehat\mu_\tau a+\widehat\nu_\tau.
$$

The theoretical guarantee is

$$
\widehat\tau_a=\tau+O_p(1/\sqrt n),
$$

or equivalently an $O_p(1/\sqrt n)$ covariance between the protected attribute and the exceedance indicator. This is groupwise quantile calibration and asymptotic balance, not exact finite-sample EQB. Its empirical application uses 198,377 newborns and reports that residual adjustment brings group-specific effective quantiles close to the nominal targets.

Causal quantile treatment-effect estimation also uses weighting to construct marginal potential-outcome distributions. For a tilting function $g(X)$, the weighted outcome distribution is estimated with weights

$$
w_{1i}=\frac{g(X_i)}{e(X_i)}Z_i,
\qquad
w_{0i}=\frac{g(X_i)}{1-e(X_i)}(1-Z_i).
$$

Weighted empirical quantiles are then compared to form a weighted QTE. IPW targets the full population when $g(X)=1$, whereas overlap weighting targets an overlap population. These procedures balance covariate distributions in population under correct propensity scores, but they do not, as presented, solve finite-sample equations that explicitly equalize covariate quantiles. They are therefore related to EQB through population balancing and weighted distribution estimation rather than strict exact calibration [2109.03757].

## 5. Distributed exact quantile computation

EQB can require exact quantile boundaries in large distributed datasets. A quick exact method for Spark uses a Greenwald–Khanna sketch to identify a near-target pivot, then performs exact selection around that pivot [2511.12025].

For a dataset of $n$ values distributed over $P$ partitions, the target is the $k$-th order statistic $x_{(k)}$. The initial GK summary provides a pivot $\pi$ whose rank is within approximately $\varepsilon n$ of the target. Each executor then counts values below, equal to, and above the pivot. If $C$ values are below $\pi$ and $E$ equal $\pi$, then $\pi$ is a valid exact answer when

$$
C<k\leq C+E.
$$

Otherwise, the residual rank is computed relative to the values above or below the pivot. Executors retain only the local candidate values that can contribute to this residual rank. Candidate sets are tree-reduced, and a final order-statistic operation returns the exact value.

The algorithm uses three logical Spark actions: construction of the approximate summary, collection of pivot counts, and tree reduction of candidates. It avoids a full-data shuffle and reports no shuffle stages. Its executor-side time matches the GK sketch asymptotically,

$$
O\left(
\frac{n}{P}\log\frac1\varepsilon
+
\frac{n}{P}\log\log\left(\varepsilon\frac nP\right)
\right),
$$

while the driver processes a candidate set whose size depends on the rank error and partition structure. The reported experiment on $10^9$ values across 120 partitions and 30 cores found approximately a $10.5$-fold improvement over Spark full sort.

For EQB applications that require $M$ equal-mass groups, exact boundary ranks can be chosen as

$$
k_j=\left\lceil\frac{jn}{M}\right\rceil,
\qquad j=1,\ldots,M-1.
$$

The resulting order-statistic values define range boundaries. Exact boundary computation does not, however, guarantee equal group cardinalities under ties. Equal-valued records may need a stable secondary key such as $(x,\mathrm{recordID})$, deterministic tie allocation, or acceptance of approximate balance at indivisible duplicate blocks. Boundary discovery and physical data placement are separate problems: GK Select addresses the former, while final range redistribution may still require substantial data movement.

## 6. EQB in distributed mixture-of-experts models

In mixture-of-experts training, EQB has a distinct operational meaning: exact global-batch quantile computation for routing margins. Let $E$ be the number of experts, $K$ the number selected per token, and

$$
\mathcal I_t=\operatorname{TopK}(z_t+b),
$$

where $z_t$ is the router-logit vector and $b$ is a token-independent expert-selection bias. For token $t$, let $\tau_t$ be the $(K+1)$-st largest adjusted logit. The margin for expert $e$ is $\tau_t-z_{t,e}$. A balanced expert should be selected for a fraction $K/E$ of tokens, so the desired bias is the corresponding empirical quantile of these margins.

If the global batch has $T$ tokens, the target rank is

$$
r=\frac{TK}{E},
$$

assuming $r$ is integral. Naive distributed QB computes rank-local quantiles and averages them, but quantiles do not commute with averaging. Histogram all-reduction is partition-invariant but approximate because the result is limited by histogram resolution.

EQB exploits BF16 representation. Each margin is encoded as an order-preserving pair of bytes,

$$
\kappa(\tau_t-z_{t,e})=(h_{t,e},\ell_{t,e}).
$$

The first all-reduce counts high-byte values and identifies the high-byte bin containing the $r$-th smallest margin. A second pass counts low-byte values only within that high-byte bin. The reconstructed value is the exact BF16 empirical quantile. The bias is then centered because adding a common constant to all expert biases does not change a top-$K$ assignment.

The communication cost is

$$
2E\cdot256
$$

32-bit integer counts per MoE layer, independent of the number of tokens. For $E=256$, the reported cost is approximately $0.5$ MiB per layer; with 48 MoE layers, it is approximately 24 MiB per optimizer step. EQB therefore avoids both rank-dependent quantile averaging and the additional binning approximation of a finite-resolution histogram [2609.28053].

The method is connected to a balanced assignment problem. Let $M_{t,e}$ indicate whether token $t$ is assigned to expert $e$. The balanced assignment set requires

$$
\sum_eM_{t,e}=K,
\qquad
\sum_tM_{t,e}=r.
$$

The balanced oracle maximizes total router score over this set. A dual representation shows that the minimizing expert bias is obtained from the $r$-th smallest collection of token-specific threshold margins. EQB computes this empirical quantile exactly, although one simultaneous finite-batch update is not necessarily a complete solution of every coupled dynamical issue in the training process.

EQB addresses global balance, not local microbatch balance. Load-Error Injection (LEI) complements it by injecting local hard-load errors into router-score gradients. In the reported 7.5-billion-parameter MoE experiments with 256 experts and top-$K$ routing with $K=6$, EQB reduced Global MaxVio from $0.92$ to $0.74$ and Local MaxVio from $5.96$ to $5.38$ relative to rank-averaged QB in 100-billion-token runs. EQB combined with normalized LEI reduced Local MaxVio further to $3.52. In 500-billion-token runs, normalized LEI reduced Local MaxVio from $5.14$ to $4.15, while Global MaxVio increased from $0.44$ to $0.63, illustrating the distinction between global and local objectives.

## 7. Guarantees, limitations, and distinctions

EQB guarantees depend on the object being balanced and on the meaning of exactness. In sensitivity analysis, exact quantile balance yields sharp identification under Tan’s marginal sensitivity model, subject to overlap, conditional-quantile regularity, and consistency of the propensity and quantile estimators. Under quantile misspecification, the estimated interval is conservative in the stated direction. In survey calibration, exactness is conditional on feasible equations, successful numerical convergence, and correctly specified target quantiles. In distributed computation, exactness is relative to the numerical representation, rank convention, duplicate policy, and percentile interpolation rule. In MoE routing, exactness is exactness of the global BF16 order statistic and does not by itself guarantee local balance or solve all training dynamics.

Several limitations recur across settings:

- **Ties and integer feasibility**: a numeric quantile boundary may not produce equal counts when many observations share the boundary value.
- **Finite versus population balance**: asymptotic or population balance does not imply exact empirical equality.
- **Selected versus entire distributions**: finitely many quantile constraints reproduce only the selected quantile points, not the full distribution.
- **Overlap and positivity**: absent support cannot be recovered through weighting alone.
- **Numerical feasibility**: too many constraints, extreme quantiles, or poor overlap can yield unstable weights or failed nonlinear equations.
- **Target-population changes**: overlap weighting improves stability but changes the estimand under heterogeneous effects.
- **Communication and placement**: exact quantile computation does not automatically provide physical redistribution into balanced partitions.
- **Global versus local balance**: global expert-routing balance does not imply microbatch-level balance.
- **Model scope**: exact relationships derived for conditional quantiles, linear quantile regression, or linear-Gaussian terminal halfspaces do not automatically extend to arbitrary distributions or dynamics.

A control-theoretic result provides another specialized quantile-balancing primitive. For a terminal Gaussian halfspace event $\{w^\top X_T\ge a\}$, changing its probability from $p_0$ to $p_1$ requires a minimum quadratic control energy

$$
E_{\min}
=
\frac{[\Phi^{-1}(p_1)-\Phi^{-1}(p_0)]^2}
{2R_T^2(w)},
$$

where

$$
R_T^2(w)=
\frac{w^\top W_c^M w}{w^\top V_T w}.
$$

The matched-filter control attains the bound. This is an exact quantile–energy equality for one terminal halfspace, not a general multi-group or distributional EQB framework [2510.17945].

Accordingly, EQB is best understood as a family of quantile-centered exactness principles rather than one algorithm. Its common structure is the replacement of unrestricted or approximate distributional control by a quantile constraint that is either solved exactly, computed exactly, or shown to characterize a sharp optimization endpoint. The implementation ranges from weighted quantile regression and calibration equations to distributed order-statistic selection, radix-based BF16 selection, and threshold biases for balanced expert assignment.

Source: https://www.emergentmind.com/topics/exact-quantile-balancing-eqb