Papers
Topics
Authors
Recent
Search
2000 character limit reached

Minimum-Ratio Estimator

Updated 15 November 2025
  • Minimum-ratio estimator is a method for estimating ratio parameters (e.g., prevalence, relative risk) with minimax optimality and controlled mean-square error.
  • It employs score functions, optimization techniques, and auxiliary information to address challenges like prior-probability shift and finite-population inference.
  • Extensions include applications in relative risk estimation and regression-ratio frameworks, ensuring unbiasedness and efficiency even under stringent error constraints.

A minimum-ratio estimator is a statistical construction designed to estimate a population parameter expressed as a ratio (e.g., prevalence, relative risk, odds ratio), with properties of minimax optimality, control over mean-square error, and—in some cases—guaranteed efficiency relative to the Cramér–Rao lower bound. These estimators appear in contexts such as quantification under prior-probability shift, finite-population inference with auxiliary variables, and design-based estimation of relative risk with controlled precision and resource allocation. The methodology is unified by the use of ratio forms and optimization over risk or mean-square error subject to specified constraints.

1. Ratio Estimators Under Prior-Probability Shift

When only a labeled source sample and an unlabeled target sample are available, quantification under prior-probability shift seeks estimation of the target prevalence πt=Ptarget(Y=1)\pi_t = P_{\mathrm{target}}(Y=1). The prior-shift assumption, Ptrain(XY)Ptarget(XY)P_{\mathrm{train}}(X|Y) \equiv P_{\mathrm{target}}(X|Y), entails that the mixture distribution of features in the target is a convex combination of class-conditional densities f0(x),f1(x)f_0(x),f_1(x) with mixture weights πt\pi_t.

The ratio estimator is constructed as: π^UR=μtμ0sμ1sμ0s,μ0s=1ns,0i:Yi=0g(Xi),    μ1s=1ns,1i:Yi=1g(Xi),    μt=1ntj=1ntg(Xj)\hat \pi_{UR} = \frac{\mu_t - \mu_0^s}{\mu_1^s - \mu_0^s}, \qquad \mu_0^s = \frac{1}{n_{s,0}} \sum_{i:Y_i=0} g(X_i),\;\; \mu_1^s = \frac{1}{n_{s,1}} \sum_{i:Y_i=1} g(X_i),\;\;\mu_t = \frac{1}{n_t} \sum_{j=1}^{n_t} g(X_j) for a scoring function gg satisfying Δ=E[g(X)Y=1]E[g(X)Y=0]0\Delta = \mathbb{E}[g(X)|Y=1]-\mathbb{E}[g(X)|Y=0] \ne 0. This estimator achieves approximate minimax optimality in 2\ell_2-risk over separable parameter classes, with the worst-case risk scaling as O(max(ns1,nt1))O(\max(n_s^{-1}, n_t^{-1})), matching the established lower bound. Trimmed forms π^R=min{1,max{0,π^UR}}\hat \pi_R = \min\{1,\max\{0,\hat \pi_{UR}\}\} are standard.

With additional knowledge of class-conditional densities, the estimator can be recast using importance weights Ptrain(XY)Ptarget(XY)P_{\mathrm{train}}(X|Y) \equiv P_{\mathrm{target}}(X|Y)0: Ptrain(XY)Ptarget(XY)P_{\mathrm{train}}(X|Y) \equiv P_{\mathrm{target}}(X|Y)1 Several classical approaches (adjusted count, EM prior-shift correction) are nested in this framework (Vaz et al., 2018).

2. Minimum-Ratio Estimation With Controlled Relative Error

In two-sample binary experiments, the minimum-ratio estimator constructs the estimate of relative risk Ptrain(XY)Ptarget(XY)P_{\mathrm{train}}(X|Y) \equiv P_{\mathrm{target}}(X|Y)2 ensuring the relative mean-square error (RMSE) is bounded by a target Ptrain(XY)Ptarget(XY)P_{\mathrm{train}}(X|Y) \equiv P_{\mathrm{target}}(X|Y)3 for all Ptrain(XY)Ptarget(XY)P_{\mathrm{train}}(X|Y) \equiv P_{\mathrm{target}}(X|Y)4, and average sample sizes are in a user-specified ratio Ptrain(XY)Ptarget(XY)P_{\mathrm{train}}(X|Y) \equiv P_{\mathrm{target}}(X|Y)5. This is achieved by a two-stage inverse binomial sampling protocol:

  • Stage I (pilot): For each population Ptrain(XY)Ptarget(XY)P_{\mathrm{train}}(X|Y) \equiv P_{\mathrm{target}}(X|Y)6, sample until Ptrain(XY)Ptarget(XY)P_{\mathrm{train}}(X|Y) \equiv P_{\mathrm{target}}(X|Y)7 successes; record number of draws Ptrain(XY)Ptarget(XY)P_{\mathrm{train}}(X|Y) \equiv P_{\mathrm{target}}(X|Y)8. Compute preliminary ratio Ptrain(XY)Ptarget(XY)P_{\mathrm{train}}(X|Y) \equiv P_{\mathrm{target}}(X|Y)9.
  • Stage II (main): Determine f0(x),f1(x)f_0(x),f_1(x)0 to satisfy

f0(x),f1(x)f_0(x),f_1(x)1

and sample until f0(x),f1(x)f_0(x),f_1(x)2 further successes (per population). The minimum-ratio (RR) estimator is then

f0(x),f1(x)f_0(x),f_1(x)3

where f0(x),f1(x)f_0(x),f_1(x)4 is the number of draws to reach f0(x),f1(x)f_0(x),f_1(x)5 successes in population f0(x),f1(x)f_0(x),f_1(x)6.

This estimator is unbiased, and achieves f0(x),f1(x)f_0(x),f_1(x)7. For small f0(x),f1(x)f_0(x),f_1(x)8, its efficiency approaches the Cramér–Rao lower bound. Batch (group) sampling operates identically, with cost adjustments for incomplete batch use (Mendo, 6 Mar 2025).

3. Ratio-Product Estimators With Auxiliary Information

For finite populations, the two-parameter ratio-product-ratio estimator uses both sample and population information from auxiliary variable f0(x),f1(x)f_0(x),f_1(x)9 to estimate the mean of study variable πt\pi_t0: πt\pi_t1 Here πt\pi_t2 are tuning parameters; πt\pi_t3 denote sample means, πt\pi_t4 the known population mean of πt\pi_t5. Specializations yield the sample mean, classical ratio, and product estimators for certain πt\pi_t6.

A first-order expansion yields bias and MSE expressions in terms of population variances πt\pi_t7, πt\pi_t8, and correlation πt\pi_t9. The MSE is minimized along the curve π^UR=μtμ0sμ1sμ0s,μ0s=1ns,0i:Yi=0g(Xi),    μ1s=1ns,1i:Yi=1g(Xi),    μt=1ntj=1ntg(Xj)\hat \pi_{UR} = \frac{\mu_t - \mu_0^s}{\mu_1^s - \mu_0^s}, \qquad \mu_0^s = \frac{1}{n_{s,0}} \sum_{i:Y_i=0} g(X_i),\;\; \mu_1^s = \frac{1}{n_{s,1}} \sum_{i:Y_i=1} g(X_i),\;\;\mu_t = \frac{1}{n_t} \sum_{j=1}^{n_t} g(X_j)0, yielding the minimum-MSE estimator: π^UR=μtμ0sμ1sμ0s,μ0s=1ns,0i:Yi=0g(Xi),    μ1s=1ns,1i:Yi=1g(Xi),    μt=1ntj=1ntg(Xj)\hat \pi_{UR} = \frac{\mu_t - \mu_0^s}{\mu_1^s - \mu_0^s}, \qquad \mu_0^s = \frac{1}{n_{s,0}} \sum_{i:Y_i=0} g(X_i),\;\; \mu_1^s = \frac{1}{n_{s,1}} \sum_{i:Y_i=1} g(X_i),\;\;\mu_t = \frac{1}{n_t} \sum_{j=1}^{n_t} g(X_j)1 For strongly correlated auxiliary information, empirical results show the minimum-ratio estimator significantly outperforms traditional approaches (Chami et al., 2012).

4. Extensions and Robustness

Multiple generalizations are available within the minimum-ratio estimator framework:

  • Combined estimator: When limited labeled target data are available, the ratio estimator can be linearly combined with plug-in estimators, using weights determined by their MSEs. This blend is provably optimal in MSE for the convex mixture (Vaz et al., 2018).
  • Regression-ratio estimator: For covariate-dependent prevalences π^UR=μtμ0sμ1sμ0s,μ0s=1ns,0i:Yi=0g(Xi),    μ1s=1ns,1i:Yi=1g(Xi),    μt=1ntj=1ntg(Xj)\hat \pi_{UR} = \frac{\mu_t - \mu_0^s}{\mu_1^s - \mu_0^s}, \qquad \mu_0^s = \frac{1}{n_{s,0}} \sum_{i:Y_i=0} g(X_i),\;\; \mu_1^s = \frac{1}{n_{s,1}} \sum_{i:Y_i=1} g(X_i),\;\;\mu_t = \frac{1}{n_t} \sum_{j=1}^{n_t} g(X_j)2, the ratio estimator extends under extra conditional independence, permitting estimation via nonparametric regression fits of π^UR=μtμ0sμ1sμ0s,μ0s=1ns,0i:Yi=0g(Xi),    μ1s=1ns,1i:Yi=1g(Xi),    μt=1ntj=1ntg(Xj)\hat \pi_{UR} = \frac{\mu_t - \mu_0^s}{\mu_1^s - \mu_0^s}, \qquad \mu_0^s = \frac{1}{n_{s,0}} \sum_{i:Y_i=0} g(X_i),\;\; \mu_1^s = \frac{1}{n_{s,1}} \sum_{i:Y_i=1} g(X_i),\;\;\mu_t = \frac{1}{n_t} \sum_{j=1}^{n_t} g(X_j)3. Consistency and convergence rates are characterized under mild regularity (Vaz et al., 2018).
  • Group sampling: In two-sample experiments, minimum-ratio sampling by group or batch is handled by simulating incremental sampling, and the performance guarantees persist with negligible overshoot variation (Mendo, 6 Mar 2025).

Crucially, weaker, empirically testable variants of the core assumptions (e.g., weak prior shift: π^UR=μtμ0sμ1sμ0s,μ0s=1ns,0i:Yi=0g(Xi),    μ1s=1ns,1i:Yi=1g(Xi),    μt=1ntj=1ntg(Xj)\hat \pi_{UR} = \frac{\mu_t - \mu_0^s}{\mu_1^s - \mu_0^s}, \qquad \mu_0^s = \frac{1}{n_{s,0}} \sum_{i:Y_i=0} g(X_i),\;\; \mu_1^s = \frac{1}{n_{s,1}} \sum_{i:Y_i=1} g(X_i),\;\;\mu_t = \frac{1}{n_t} \sum_{j=1}^{n_t} g(X_j)4) suffice for consistency and confidence interval construction. Verification procedures based on the convex hull relationship of CDFs support practical deployment.

5. Optimality and Performance

For quantification, the minimax lower bound for mean squared error cannot improve upon π^UR=μtμ0sμ1sμ0s,μ0s=1ns,0i:Yi=0g(Xi),    μ1s=1ns,1i:Yi=1g(Xi),    μt=1ntj=1ntg(Xj)\hat \pi_{UR} = \frac{\mu_t - \mu_0^s}{\mu_1^s - \mu_0^s}, \qquad \mu_0^s = \frac{1}{n_{s,0}} \sum_{i:Y_i=0} g(X_i),\;\; \mu_1^s = \frac{1}{n_{s,1}} \sum_{i:Y_i=1} g(X_i),\;\;\mu_t = \frac{1}{n_t} \sum_{j=1}^{n_t} g(X_j)5, as established by Bayes and Le Cam type arguments. The ratio estimator matches this rate, and hence is approximately minimax. For relative risk, the minimum-ratio estimator is unbiased and has relative mean-square error strictly bounded by π^UR=μtμ0sμ1sμ0s,μ0s=1ns,0i:Yi=0g(Xi),    μ1s=1ns,1i:Yi=1g(Xi),    μt=1ntj=1ntg(Xj)\hat \pi_{UR} = \frac{\mu_t - \mu_0^s}{\mu_1^s - \mu_0^s}, \qquad \mu_0^s = \frac{1}{n_{s,0}} \sum_{i:Y_i=0} g(X_i),\;\; \mu_1^s = \frac{1}{n_{s,1}} \sum_{i:Y_i=1} g(X_i),\;\;\mu_t = \frac{1}{n_t} \sum_{j=1}^{n_t} g(X_j)6 everywhere; for small error requirements, its efficiency (variance relative to the Cramér–Rao lower bound) converges to unity (Mendo, 6 Mar 2025).

For auxiliary-based population mean estimation, the minimum-MSE estimator achieves the classical regression-type optimum, with the empirical MSE and coverage of confidence intervals demonstrating substantial relative efficiency gains over sample mean, ratio, or product estimators (Chami et al., 2012).

6. Algorithmic Construction and Implementation

The minimum-ratio estimator methodology is characterized by systematic algebraic derivations determining parameters—either tuning π^UR=μtμ0sμ1sμ0s,μ0s=1ns,0i:Yi=0g(Xi),    μ1s=1ns,1i:Yi=1g(Xi),    μt=1ntj=1ntg(Xj)\hat \pi_{UR} = \frac{\mu_t - \mu_0^s}{\mu_1^s - \mu_0^s}, \qquad \mu_0^s = \frac{1}{n_{s,0}} \sum_{i:Y_i=0} g(X_i),\;\; \mu_1^s = \frac{1}{n_{s,1}} \sum_{i:Y_i=1} g(X_i),\;\;\mu_t = \frac{1}{n_t} \sum_{j=1}^{n_t} g(X_j)7, or sequential stopping rules π^UR=μtμ0sμ1sμ0s,μ0s=1ns,0i:Yi=0g(Xi),    μ1s=1ns,1i:Yi=1g(Xi),    μt=1ntj=1ntg(Xj)\hat \pi_{UR} = \frac{\mu_t - \mu_0^s}{\mu_1^s - \mu_0^s}, \qquad \mu_0^s = \frac{1}{n_{s,0}} \sum_{i:Y_i=0} g(X_i),\;\; \mu_1^s = \frac{1}{n_{s,1}} \sum_{i:Y_i=1} g(X_i),\;\;\mu_t = \frac{1}{n_t} \sum_{j=1}^{n_t} g(X_j)8—to meet targeted precision. For relative-risk, the construction proceeds via:

  1. Solving for pilot-stage sizes (π^UR=μtμ0sμ1sμ0s,μ0s=1ns,0i:Yi=0g(Xi),    μ1s=1ns,1i:Yi=1g(Xi),    μt=1ntj=1ntg(Xj)\hat \pi_{UR} = \frac{\mu_t - \mu_0^s}{\mu_1^s - \mu_0^s}, \qquad \mu_0^s = \frac{1}{n_{s,0}} \sum_{i:Y_i=0} g(X_i),\;\; \mu_1^s = \frac{1}{n_{s,1}} \sum_{i:Y_i=1} g(X_i),\;\;\mu_t = \frac{1}{n_t} \sum_{j=1}^{n_t} g(X_j)9) from curvature conditions.
  2. Computing gg0 to exactly achieve the mean-square error and allocation targets.
  3. Executing sampling and estimator calculations accordingly.

For quantification, the workflow is:

  • Compute class means of the score function gg1 over labeled and target samples.
  • Form the ratio of means estimator (possibly trimmed).
  • Combine with plug-in estimators as needed or incorporate regression for prevalence as function of covariates.

The explicit forms support straightforward algorithmic implementation, including for large-scale or streaming contexts, and the underlying variance control properties remain valid under a wide range of scenarios due to their non-reliance on asymptotics or strong distributional assumptions.

7. Applications and Empirical Performance

Minimum-ratio estimators are applied in diverse contexts, such as:

  • Quantification of prevalence in target populations under distribution shift (e.g., sentiment analysis domain adaptation) (Vaz et al., 2018).
  • Design-based survey estimation with auxiliary variables, especially when high correlation allows for dramatic MSE reduction (Chami et al., 2012).
  • Clinical and epidemiological studies requiring strict control on the error in estimation of relative risk, regardless of underlying proportions (Mendo, 6 Mar 2025).

Empirical validation demonstrates substantial improvements in MSE and coverage, robustness to misspecification of underlying distributions (given testable assumptions), and scalability to high-throughput sampling schemes.


The development and analysis of minimum-ratio estimators unify minimax, unbiasedness, and efficiency properties across quantification, finite-population estimation, and two-sample problems. Their construction enables both rigorous statistical guarantees and practical deployability in heterogeneous data regimes.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Minimum-ratio Estimator.