Heteroscedasticity-Adjusted Ranking (HART)
- HART is a statistical framework for multiple testing that accounts for heteroscedasticity by using both summary statistics and their associated variances.
- It employs a conditional local false discovery rate to rank hypotheses, enabling more effective thresholding than standard z-score based methods.
- Empirical studies show that HART increases detection power in large, heterogeneous datasets by leveraging variance information ignored in conventional approaches.
HART, short for Heterocedasticity-Adjusted Ranking and Thresholding, is a two-step multiple testing framework for large-scale heterogeneous data that bypasses standardization and uses both a summary statistic and its unit-specific variance for ranking and thresholding. Its central claim is that standardization to -scores or -values can discard information carried by the variance structure of the alternative distribution, especially under heteroscedasticity. The methodology, theory, and empirical studies were developed by Fu, Gang, James, and Sun in “Heterocedasticity-Adjusted Ranking and Thresholding for Large-Scale Multiple Testing” (Fu et al., 2019).
1. Conceptual basis and statistical setting
HART is designed for settings in which, for , one observes a summary statistic and a variance or standard deviation or , with heteroscedasticity across units. The inferential task is
with signal indicator following a model (Fu et al., 2019).
The canonical probabilistic model is Gaussian with heteroscedastic errors:
where 0 is a point mass at 1. In much of the analysis, 2 are treated as known or consistently estimated (Fu et al., 2019).
Conditioning on 3, the marginal mixture for 4 is
5
with
6
The motivating observation is that standardization to 7 maps all nulls to 8, but also transforms the alternative to
9
This subsumes the variance structure of the alternative. Consequently, for the same 0-value, small-1 units are more consistent with the alternative than large-2 units. HART exploits this by conditioning on 3 in both ranking and thresholding (Fu et al., 2019).
2. Oracle ranking and thresholding rule
The core HART ranking statistic is the local false discovery rate conditional on 4:
5
This is a Bayesian likelihood-ratio-based score; smaller 6 indicates stronger evidence against the null (Fu et al., 2019).
Let 7 denote the marginal false discovery rate level, 8, of the rule that rejects all 9 with 0. The oracle HART rule is
1
This ranking-plus-threshold construction uses both 2 and 3 explicitly and is optimal within the class of rules based on 4 (Fu et al., 2019).
The contrast with standardized procedures is explicit. For standardized values 5, the classical local false discovery rate is
6
and the oracle 7-value rule is 8. The oracle 9-value rule is 0. These procedures do not condition on 1 and therefore discard heteroscedastic information (Fu et al., 2019).
A key structural special case is the homoscedastic regime 2. Then 3 is a monotone transform of 4, and HART reduces to the standard local false discovery rate procedure (Fu et al., 2019).
3. Error control, power, and asymptotic theory
HART is developed for false discovery rate control while maximizing expected true positives. The paper defines
5
where
6
Power is measured by
7
The oracle optimality result states that if 8 is the collection of all procedures based on 9 with 0, then
1
In particular,
2
The intuition given in the paper is that, conditional on 3, thresholding 4 is a monotone likelihood ratio decision under the mixture model, so its rejection region maximizes alternative mass at fixed expected false discoveries (Fu et al., 2019).
For the data-driven procedure, HART estimates both 5 and 6. Its asymptotic analysis assumes four conditions. First, bounded supports:
7
Second, regularity of the kernel 8: positive, bounded, symmetric, with
9
and 0 having bounded, continuous second derivative and being square-integrable. Third, bandwidth conditions:
1
2
Fourth,
3
Under these conditions, the consistency proposition states
4
The asymptotic validity and optimality theorem further gives
5
and
6
where 7 denotes the data-driven HART rule (Fu et al., 2019).
4. Estimation strategy and computational construction
HART estimates the conditional mixture density through a stabilized decomposition,
8
where 9 is a weighted bivariate kernel estimator with variable bandwidth in 0 to account for heteroscedasticity:
1
The weights 2 approximate 3 and are refined in two steps (Fu et al., 2019).
The initialization uses
4
with
5
The updated weights are
6
The final score is
7
Thresholding is empirical. After sorting
8
HART selects
9
and rejects 0 (Fu et al., 2019).
The naive computational cost of evaluating 1 for all 2 is 3. The paper notes that fast kernel summation, including KD-trees, FFT-based approximations, sparsity via compact kernels, batching, or GPUs, can reduce wall time. A jackknife variant, implemented by leaving one point out in the density estimate, is described as stabilizing estimates with negligible asymptotic cost (Fu et al., 2019).
The base model assumes 4 independent of 5, but the paper also gives a structured extension for dependence:
6
with
7
Misspecifying this dependence may reduce power, while the basic HART machinery still yields valid ranking under mild conditions (Fu et al., 2019).
5. Empirical behavior and applications
The paper’s finite-mixture illustration uses 8, 9, 0, and 1. The average power of three oracle procedures is reported as
2
The full-data oracle therefore more than doubles the average power relative to the 3-value oracle (Fu et al., 2019).
In continuous-heteroscedastic simulations,
4
with 5 and 6, all methods controlled FDR at nominal levels, while data-driven HART was slightly conservative at times. The reported power pattern is that data-driven HART uniformly exceeded BH and the 7-oracle proxy, often closing much of the gap to the full-data oracle; gains increased with stronger heterogeneity and sparser signals, with relative power improvements frequently in the 20–60% range versus BH and ZOR at matched FDR (Fu et al., 2019).
In the two-group heteroscedastic setting, with 8 equiprobable, 9, 00, 01, and 02, the data-driven procedure approximately matched the oracle in both FDR and power. HART applied differing 03-cutoffs per group, increasing the cutoff for the large-variance group, and the power gap relative to ZOR grew with 04 (Fu et al., 2019).
The principal empirical case study is a microarray analysis of myeloma. The dataset contained 12,625 genes; 36 myeloma patients had MRI-detected focal lesions and 137 had no lesions. Gene-wise contrasts 05 and standard errors 06 were computed. Because the theoretical 07 null was too narrow, BH and AZ used the empirical null 08. For stable HART density estimation, genes with 09 were analyzed, leaving 12,172 genes. At 10, BH discovered 8 genes (0.07%), AZ 25 (0.2%), and HART 122 (1%). The rejection boundaries for BH and AZ depended only on 11, whereas HART’s decisions varied with both 12 and 13, favoring small-14 genes and avoiding large-15 genes (Fu et al., 2019).
6. Relation to other procedures, practical regime, and limitations
HART is explicitly positioned against several established multiple-testing procedures. BH on standardized 16-values or 17-scores ignores structure in the alternative and often uses symmetric two-sided rejection. Weighted BH, IHW, and AdaPT use covariates to pre-order or weight 18-values, but naïvely ordering by 19 can be anti-informative because correctness depends on the relationship between 20 and signal prevalence or strength. HART instead performs a semiparametric empirical Bayes integration of 21 through conditional density modeling (Fu et al., 2019).
In the special case where 22 takes one of a few discrete values, HART reduces to conditional local false discovery rate analyzed separately per group, with group-specific thresholds. Relative to classical local false discovery rate procedures that estimate 23, HART estimates 24 and 25 conditionally, restoring the alternative variance structure lost after standardization (Fu et al., 2019).
The paper identifies the regime in which HART is most beneficial: clear heteroscedasticity, effect sizes whose discriminability depends on 26, and large 27, typically in the thousands or more, so that conditional density estimation is accurate. The recommended diagnostics are a histogram of 28, a scatter of 29 versus 30, empirical 31 plotted against 32 by 33 strata, and direct comparison of HART and 34-based rejection boundaries. For tuning, the paper recommends estimating 35 by methods robust to alternatives such as Jin–Cai, choosing bandwidths by rules of thumb on a subset enriched for signals, and considering jackknifed density estimation. For unknown 36, it recommends consistent variance estimators and, if needed, 37-approximations for 38 (Fu et al., 2019).
The limitations are equally explicit. HART requires large 39 for stable bivariate density estimation and is computationally heavier than BH or ZOR, although it may be faster than some covariate-adaptive procedures such as AdaPT in many settings. Its theoretical development assumes independence or weak dependence and known null scale per 40. The extension of empirical-null modeling under heteroscedasticity is nontrivial when the theoretical null is inadequate. Finally, misspecification of 41–42 dependence can reduce power, motivating structured extensions with 43 and 44 (Fu et al., 2019).
HART’s defining methodological claim is therefore narrow but consequential: when unit-specific variances are themselves informative under the alternative, ranking by
45
and thresholding to control FDR can strictly dominate procedures based only on standardized statistics. In the framework developed in (Fu et al., 2019), this gain is formalized by oracle optimality, asymptotic validity, and repeated empirical power improvements at fixed FDR.