Papers
Topics
Authors
Recent
Search
2000 character limit reached

Heteroscedasticity-Adjusted Ranking (HART)

Updated 14 July 2026
  • HART is a statistical framework for multiple testing that accounts for heteroscedasticity by using both summary statistics and their associated variances.
  • It employs a conditional local false discovery rate to rank hypotheses, enabling more effective thresholding than standard z-score based methods.
  • Empirical studies show that HART increases detection power in large, heterogeneous datasets by leveraging variance information ignored in conventional approaches.

HART, short for Heterocedasticity-Adjusted Ranking and Thresholding, is a two-step multiple testing framework for large-scale heterogeneous data that bypasses standardization and uses both a summary statistic and its unit-specific variance for ranking and thresholding. Its central claim is that standardization to zz-scores or pp-values can discard information carried by the variance structure of the alternative distribution, especially under heteroscedasticity. The methodology, theory, and empirical studies were developed by Fu, Gang, James, and Sun in “Heterocedasticity-Adjusted Ranking and Thresholding for Large-Scale Multiple Testing” (Fu et al., 2019).

1. Conceptual basis and statistical setting

HART is designed for settings in which, for i=1,,mi=1,\dots,m, one observes a summary statistic XiX_i and a variance or standard deviation si2s_i^2 or sis_i, with heteroscedasticity across units. The inferential task is

H0,i:μi=0versusH1,i:μi0,H_{0,i}:\mu_i=0 \qquad \text{versus} \qquad H_{1,i}:\mu_i\neq 0,

with signal indicator θi=1(μi0)\theta_i=1(\mu_i\neq 0) following a Bernoulli(π)\mathrm{Bernoulli}(\pi) model (Fu et al., 2019).

The canonical probabilistic model is Gaussian with heteroscedastic errors:

Xiμi,σi2indN(μi,σi2),μiiid(1π)δ0+πgμ,σi2iidgσ,X_i \mid \mu_i,\sigma_i^2 \overset{ind}{\sim} N(\mu_i,\sigma_i^2),\qquad \mu_i \overset{iid}{\sim} (1-\pi)\delta_0+\pi g_\mu,\qquad \sigma_i^2 \overset{iid}{\sim} g_\sigma,

where pp0 is a point mass at pp1. In much of the analysis, pp2 are treated as known or consistently estimated (Fu et al., 2019).

Conditioning on pp3, the marginal mixture for pp4 is

pp5

with

pp6

The motivating observation is that standardization to pp7 maps all nulls to pp8, but also transforms the alternative to

pp9

This subsumes the variance structure of the alternative. Consequently, for the same i=1,,mi=1,\dots,m0-value, small-i=1,,mi=1,\dots,m1 units are more consistent with the alternative than large-i=1,,mi=1,\dots,m2 units. HART exploits this by conditioning on i=1,,mi=1,\dots,m3 in both ranking and thresholding (Fu et al., 2019).

2. Oracle ranking and thresholding rule

The core HART ranking statistic is the local false discovery rate conditional on i=1,,mi=1,\dots,m4:

i=1,,mi=1,\dots,m5

This is a Bayesian likelihood-ratio-based score; smaller i=1,,mi=1,\dots,m6 indicates stronger evidence against the null (Fu et al., 2019).

Let i=1,,mi=1,\dots,m7 denote the marginal false discovery rate level, i=1,,mi=1,\dots,m8, of the rule that rejects all i=1,,mi=1,\dots,m9 with XiX_i0. The oracle HART rule is

XiX_i1

This ranking-plus-threshold construction uses both XiX_i2 and XiX_i3 explicitly and is optimal within the class of rules based on XiX_i4 (Fu et al., 2019).

The contrast with standardized procedures is explicit. For standardized values XiX_i5, the classical local false discovery rate is

XiX_i6

and the oracle XiX_i7-value rule is XiX_i8. The oracle XiX_i9-value rule is si2s_i^20. These procedures do not condition on si2s_i^21 and therefore discard heteroscedastic information (Fu et al., 2019).

A key structural special case is the homoscedastic regime si2s_i^22. Then si2s_i^23 is a monotone transform of si2s_i^24, and HART reduces to the standard local false discovery rate procedure (Fu et al., 2019).

3. Error control, power, and asymptotic theory

HART is developed for false discovery rate control while maximizing expected true positives. The paper defines

si2s_i^25

where

si2s_i^26

Power is measured by

si2s_i^27

The oracle optimality result states that if si2s_i^28 is the collection of all procedures based on si2s_i^29 with sis_i0, then

sis_i1

In particular,

sis_i2

The intuition given in the paper is that, conditional on sis_i3, thresholding sis_i4 is a monotone likelihood ratio decision under the mixture model, so its rejection region maximizes alternative mass at fixed expected false discoveries (Fu et al., 2019).

For the data-driven procedure, HART estimates both sis_i5 and sis_i6. Its asymptotic analysis assumes four conditions. First, bounded supports:

sis_i7

Second, regularity of the kernel sis_i8: positive, bounded, symmetric, with

sis_i9

and H0,i:μi=0versusH1,i:μi0,H_{0,i}:\mu_i=0 \qquad \text{versus} \qquad H_{1,i}:\mu_i\neq 0,0 having bounded, continuous second derivative and being square-integrable. Third, bandwidth conditions:

H0,i:μi=0versusH1,i:μi0,H_{0,i}:\mu_i=0 \qquad \text{versus} \qquad H_{1,i}:\mu_i\neq 0,1

H0,i:μi=0versusH1,i:μi0,H_{0,i}:\mu_i=0 \qquad \text{versus} \qquad H_{1,i}:\mu_i\neq 0,2

Fourth,

H0,i:μi=0versusH1,i:μi0,H_{0,i}:\mu_i=0 \qquad \text{versus} \qquad H_{1,i}:\mu_i\neq 0,3

Under these conditions, the consistency proposition states

H0,i:μi=0versusH1,i:μi0,H_{0,i}:\mu_i=0 \qquad \text{versus} \qquad H_{1,i}:\mu_i\neq 0,4

The asymptotic validity and optimality theorem further gives

H0,i:μi=0versusH1,i:μi0,H_{0,i}:\mu_i=0 \qquad \text{versus} \qquad H_{1,i}:\mu_i\neq 0,5

and

H0,i:μi=0versusH1,i:μi0,H_{0,i}:\mu_i=0 \qquad \text{versus} \qquad H_{1,i}:\mu_i\neq 0,6

where H0,i:μi=0versusH1,i:μi0,H_{0,i}:\mu_i=0 \qquad \text{versus} \qquad H_{1,i}:\mu_i\neq 0,7 denotes the data-driven HART rule (Fu et al., 2019).

4. Estimation strategy and computational construction

HART estimates the conditional mixture density through a stabilized decomposition,

H0,i:μi=0versusH1,i:μi0,H_{0,i}:\mu_i=0 \qquad \text{versus} \qquad H_{1,i}:\mu_i\neq 0,8

where H0,i:μi=0versusH1,i:μi0,H_{0,i}:\mu_i=0 \qquad \text{versus} \qquad H_{1,i}:\mu_i\neq 0,9 is a weighted bivariate kernel estimator with variable bandwidth in θi=1(μi0)\theta_i=1(\mu_i\neq 0)0 to account for heteroscedasticity:

θi=1(μi0)\theta_i=1(\mu_i\neq 0)1

The weights θi=1(μi0)\theta_i=1(\mu_i\neq 0)2 approximate θi=1(μi0)\theta_i=1(\mu_i\neq 0)3 and are refined in two steps (Fu et al., 2019).

The initialization uses

θi=1(μi0)\theta_i=1(\mu_i\neq 0)4

with

θi=1(μi0)\theta_i=1(\mu_i\neq 0)5

The updated weights are

θi=1(μi0)\theta_i=1(\mu_i\neq 0)6

The final score is

θi=1(μi0)\theta_i=1(\mu_i\neq 0)7

Thresholding is empirical. After sorting

θi=1(μi0)\theta_i=1(\mu_i\neq 0)8

HART selects

θi=1(μi0)\theta_i=1(\mu_i\neq 0)9

and rejects Bernoulli(π)\mathrm{Bernoulli}(\pi)0 (Fu et al., 2019).

The naive computational cost of evaluating Bernoulli(π)\mathrm{Bernoulli}(\pi)1 for all Bernoulli(π)\mathrm{Bernoulli}(\pi)2 is Bernoulli(π)\mathrm{Bernoulli}(\pi)3. The paper notes that fast kernel summation, including KD-trees, FFT-based approximations, sparsity via compact kernels, batching, or GPUs, can reduce wall time. A jackknife variant, implemented by leaving one point out in the density estimate, is described as stabilizing estimates with negligible asymptotic cost (Fu et al., 2019).

The base model assumes Bernoulli(π)\mathrm{Bernoulli}(\pi)4 independent of Bernoulli(π)\mathrm{Bernoulli}(\pi)5, but the paper also gives a structured extension for dependence:

Bernoulli(π)\mathrm{Bernoulli}(\pi)6

with

Bernoulli(π)\mathrm{Bernoulli}(\pi)7

Misspecifying this dependence may reduce power, while the basic HART machinery still yields valid ranking under mild conditions (Fu et al., 2019).

5. Empirical behavior and applications

The paper’s finite-mixture illustration uses Bernoulli(π)\mathrm{Bernoulli}(\pi)8, Bernoulli(π)\mathrm{Bernoulli}(\pi)9, Xiμi,σi2indN(μi,σi2),μiiid(1π)δ0+πgμ,σi2iidgσ,X_i \mid \mu_i,\sigma_i^2 \overset{ind}{\sim} N(\mu_i,\sigma_i^2),\qquad \mu_i \overset{iid}{\sim} (1-\pi)\delta_0+\pi g_\mu,\qquad \sigma_i^2 \overset{iid}{\sim} g_\sigma,0, and Xiμi,σi2indN(μi,σi2),μiiid(1π)δ0+πgμ,σi2iidgσ,X_i \mid \mu_i,\sigma_i^2 \overset{ind}{\sim} N(\mu_i,\sigma_i^2),\qquad \mu_i \overset{iid}{\sim} (1-\pi)\delta_0+\pi g_\mu,\qquad \sigma_i^2 \overset{iid}{\sim} g_\sigma,1. The average power of three oracle procedures is reported as

Xiμi,σi2indN(μi,σi2),μiiid(1π)δ0+πgμ,σi2iidgσ,X_i \mid \mu_i,\sigma_i^2 \overset{ind}{\sim} N(\mu_i,\sigma_i^2),\qquad \mu_i \overset{iid}{\sim} (1-\pi)\delta_0+\pi g_\mu,\qquad \sigma_i^2 \overset{iid}{\sim} g_\sigma,2

The full-data oracle therefore more than doubles the average power relative to the Xiμi,σi2indN(μi,σi2),μiiid(1π)δ0+πgμ,σi2iidgσ,X_i \mid \mu_i,\sigma_i^2 \overset{ind}{\sim} N(\mu_i,\sigma_i^2),\qquad \mu_i \overset{iid}{\sim} (1-\pi)\delta_0+\pi g_\mu,\qquad \sigma_i^2 \overset{iid}{\sim} g_\sigma,3-value oracle (Fu et al., 2019).

In continuous-heteroscedastic simulations,

Xiμi,σi2indN(μi,σi2),μiiid(1π)δ0+πgμ,σi2iidgσ,X_i \mid \mu_i,\sigma_i^2 \overset{ind}{\sim} N(\mu_i,\sigma_i^2),\qquad \mu_i \overset{iid}{\sim} (1-\pi)\delta_0+\pi g_\mu,\qquad \sigma_i^2 \overset{iid}{\sim} g_\sigma,4

with Xiμi,σi2indN(μi,σi2),μiiid(1π)δ0+πgμ,σi2iidgσ,X_i \mid \mu_i,\sigma_i^2 \overset{ind}{\sim} N(\mu_i,\sigma_i^2),\qquad \mu_i \overset{iid}{\sim} (1-\pi)\delta_0+\pi g_\mu,\qquad \sigma_i^2 \overset{iid}{\sim} g_\sigma,5 and Xiμi,σi2indN(μi,σi2),μiiid(1π)δ0+πgμ,σi2iidgσ,X_i \mid \mu_i,\sigma_i^2 \overset{ind}{\sim} N(\mu_i,\sigma_i^2),\qquad \mu_i \overset{iid}{\sim} (1-\pi)\delta_0+\pi g_\mu,\qquad \sigma_i^2 \overset{iid}{\sim} g_\sigma,6, all methods controlled FDR at nominal levels, while data-driven HART was slightly conservative at times. The reported power pattern is that data-driven HART uniformly exceeded BH and the Xiμi,σi2indN(μi,σi2),μiiid(1π)δ0+πgμ,σi2iidgσ,X_i \mid \mu_i,\sigma_i^2 \overset{ind}{\sim} N(\mu_i,\sigma_i^2),\qquad \mu_i \overset{iid}{\sim} (1-\pi)\delta_0+\pi g_\mu,\qquad \sigma_i^2 \overset{iid}{\sim} g_\sigma,7-oracle proxy, often closing much of the gap to the full-data oracle; gains increased with stronger heterogeneity and sparser signals, with relative power improvements frequently in the 20–60% range versus BH and ZOR at matched FDR (Fu et al., 2019).

In the two-group heteroscedastic setting, with Xiμi,σi2indN(μi,σi2),μiiid(1π)δ0+πgμ,σi2iidgσ,X_i \mid \mu_i,\sigma_i^2 \overset{ind}{\sim} N(\mu_i,\sigma_i^2),\qquad \mu_i \overset{iid}{\sim} (1-\pi)\delta_0+\pi g_\mu,\qquad \sigma_i^2 \overset{iid}{\sim} g_\sigma,8 equiprobable, Xiμi,σi2indN(μi,σi2),μiiid(1π)δ0+πgμ,σi2iidgσ,X_i \mid \mu_i,\sigma_i^2 \overset{ind}{\sim} N(\mu_i,\sigma_i^2),\qquad \mu_i \overset{iid}{\sim} (1-\pi)\delta_0+\pi g_\mu,\qquad \sigma_i^2 \overset{iid}{\sim} g_\sigma,9, pp00, pp01, and pp02, the data-driven procedure approximately matched the oracle in both FDR and power. HART applied differing pp03-cutoffs per group, increasing the cutoff for the large-variance group, and the power gap relative to ZOR grew with pp04 (Fu et al., 2019).

The principal empirical case study is a microarray analysis of myeloma. The dataset contained 12,625 genes; 36 myeloma patients had MRI-detected focal lesions and 137 had no lesions. Gene-wise contrasts pp05 and standard errors pp06 were computed. Because the theoretical pp07 null was too narrow, BH and AZ used the empirical null pp08. For stable HART density estimation, genes with pp09 were analyzed, leaving 12,172 genes. At pp10, BH discovered 8 genes (0.07%), AZ 25 (0.2%), and HART 122 (1%). The rejection boundaries for BH and AZ depended only on pp11, whereas HART’s decisions varied with both pp12 and pp13, favoring small-pp14 genes and avoiding large-pp15 genes (Fu et al., 2019).

6. Relation to other procedures, practical regime, and limitations

HART is explicitly positioned against several established multiple-testing procedures. BH on standardized pp16-values or pp17-scores ignores structure in the alternative and often uses symmetric two-sided rejection. Weighted BH, IHW, and AdaPT use covariates to pre-order or weight pp18-values, but naïvely ordering by pp19 can be anti-informative because correctness depends on the relationship between pp20 and signal prevalence or strength. HART instead performs a semiparametric empirical Bayes integration of pp21 through conditional density modeling (Fu et al., 2019).

In the special case where pp22 takes one of a few discrete values, HART reduces to conditional local false discovery rate analyzed separately per group, with group-specific thresholds. Relative to classical local false discovery rate procedures that estimate pp23, HART estimates pp24 and pp25 conditionally, restoring the alternative variance structure lost after standardization (Fu et al., 2019).

The paper identifies the regime in which HART is most beneficial: clear heteroscedasticity, effect sizes whose discriminability depends on pp26, and large pp27, typically in the thousands or more, so that conditional density estimation is accurate. The recommended diagnostics are a histogram of pp28, a scatter of pp29 versus pp30, empirical pp31 plotted against pp32 by pp33 strata, and direct comparison of HART and pp34-based rejection boundaries. For tuning, the paper recommends estimating pp35 by methods robust to alternatives such as Jin–Cai, choosing bandwidths by rules of thumb on a subset enriched for signals, and considering jackknifed density estimation. For unknown pp36, it recommends consistent variance estimators and, if needed, pp37-approximations for pp38 (Fu et al., 2019).

The limitations are equally explicit. HART requires large pp39 for stable bivariate density estimation and is computationally heavier than BH or ZOR, although it may be faster than some covariate-adaptive procedures such as AdaPT in many settings. Its theoretical development assumes independence or weak dependence and known null scale per pp40. The extension of empirical-null modeling under heteroscedasticity is nontrivial when the theoretical null is inadequate. Finally, misspecification of pp41–pp42 dependence can reduce power, motivating structured extensions with pp43 and pp44 (Fu et al., 2019).

HART’s defining methodological claim is therefore narrow but consequential: when unit-specific variances are themselves informative under the alternative, ranking by

pp45

and thresholding to control FDR can strictly dominate procedures based only on standardized statistics. In the framework developed in (Fu et al., 2019), this gain is formalized by oracle optimality, asymptotic validity, and repeated empirical power improvements at fixed FDR.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to HART.