Papers
Topics
Authors
Recent
Search
2000 character limit reached

AdaFilter-Bon: Bonferroni Adaptive Filtering

Updated 9 July 2026
  • AdaFilter-Bon is a partial conjunction testing method that utilizes a Bonferroni-based p-value combination to detect replicability across multiple studies.
  • It adaptively filters out unlikely candidates before applying the rejection rule, thereby reducing the conservativeness of direct multiple-testing correction.
  • Empirical evaluations demonstrate that AdaFilter-Bon improves power while controlling familywise error rate and per-family error rate under independence assumptions.

Searching arXiv for recent and foundational papers on AdaFilter-Bon and related terminology. arxiv_search(query="AdaFilter-Bon partial conjunction adaptive filtering best-of-n", max_results=10) AdaFilter-Bon is the Bonferroni-based AdaFilter procedure for familywise error rate control in partial conjunction testing, designed to detect signals that replicate across multiple studies while avoiding the severe conservativeness of direct multiple-testing correction on partial conjunction pp-values. In the formulation given in “Detecting Multiple Replicating Signals using Adaptive Filtering Procedures” (Wang et al., 2016), the method increases power by adaptively filtering out unlikely candidates of partial conjunction nulls, then applying a Bonferroni-style rejection rule on the filtered set. Later work generalized the same filtering logic to kk-family-wise error rate control and analyzed how post-filter null proportion estimation can reduce conservativeness (Tran, 21 Aug 2025).

1. Problem setting and statistical objective

AdaFilter-Bon is defined in the setting of replicability analysis across KK studies and MM hypotheses or features. For hypothesis j=1,,Mj=1,\dots,M and study i=1,,Ki=1,\dots,K, let pijp_{ij} be the valid per-study pp-value, and let the ordered values be P(1)jP(K)jP_{(1)j} \le \cdots \le P_{(K)j} (Wang et al., 2016).

The target inferential object is the partial conjunction null

H0jr/K:fewer than r of the K base hypotheses are non-null.H_{0j}^{r/K}:\quad \text{fewer than } r \text{ of the } K \text{ base hypotheses are non-null}.

Rejecting kk0 asserts replicability in at least kk1 studies. This distinguishes replicability from ordinary meta-analysis: the goal is not merely to aggregate evidence across studies, but specifically to identify signals that are discovered in multiple studies (Wang et al., 2016).

For simultaneous testing, let kk2 if kk3 is rejected and kk4 otherwise, with total discoveries kk5 and false discoveries

kk6

The principal error criteria are

kk7

AdaFilter-Bon targets FWER and, more strongly, PFER (Wang et al., 2016).

2. Partial conjunction kk8-values and the role of Bonferroni combination

For a single hypothesis with per-study kk9-values KK0, the framework admits several valid partial conjunction KK1-values under independence. The Bonferroni-based partial conjunction KK2-value is

KK3

The Simes-type partial conjunction KK4-value is

KK5

The Fisher-type partial conjunction KK6-value is

KK7

These constructions apply only to the largest KK8 KK9-values because the least favorable null configurations are those with exactly MM0 non-nulls. For MM1, all three reduce to MM2 (Wang et al., 2016).

AdaFilter-Bon uses the Bonferroni partial conjunction statistic internally. This is not an incidental design choice: the filtering argument and the resulting conditional validity inequality are stated in terms of the Bonferroni-based order-statistic construction (Wang et al., 2016). A later MM3-FWER treatment preserved the same choice, writing for feature MM4 and replication level MM5,

MM6

with MM7 by construction (Tran, 21 Aug 2025).

3. Core construction of AdaFilter-Bon

AdaFilter-Bon attaches two statistics to each hypothesis MM8. The filter statistic is

MM9

and the selection statistic is

j=1,,Mj=1,\dots,M0

The method’s key intuition is that hypotheses with large j=1,,Mj=1,\dots,M1 are unlikely to be least favorable partial conjunction nulls, so they can be filtered before multiplicity correction (Wang et al., 2016).

The central validity statement is the conditional validity lemma. Under independence across studies, if j=1,,Mj=1,\dots,M2 is true then for any fixed j=1,,Mj=1,\dots,M3,

j=1,,Mj=1,\dots,M4

Equivalently,

j=1,,Mj=1,\dots,M5

because j=1,,Mj=1,\dots,M6 always holds. This inequality is the technical device that makes adaptive thresholding valid (Wang et al., 2016).

The Bonferroni AdaFilter threshold is

j=1,,Mj=1,\dots,M7

The rejection rule is

j=1,,Mj=1,\dots,M8

An equivalent adjusted-j=1,,Mj=1,\dots,M9-value formulation ranks the selection statistics i=1,,Ki=1,\dots,K0 and defines

i=1,,Ki=1,\dots,K1

i=1,,Ki=1,\dots,K2

Reject i=1,,Ki=1,\dots,K3 if

i=1,,Ki=1,\dots,K4

This produces the same rejections as the i=1,,Ki=1,\dots,K5-based rule and is the computationally convenient formulation emphasized in the procedure description (Wang et al., 2016).

4. Error control, monotonicity, and why power increases

The assumptions for finite-i=1,,Ki=1,\dots,K6 FWER and PFER control are explicit. The base i=1,,Ki=1,\dots,K7-values must be valid, independence across studies is required for the conditional validity lemma, and the proof of finite-i=1,,Ki=1,\dots,K8 FWER/PFER control uses independence of all i=1,,Ki=1,\dots,K9 base pijp_{ij}0-values (Wang et al., 2016).

Under those assumptions, AdaFilter-Bon satisfies

pijp_{ij}1

The proof strategy introduces a data-dependent auxiliary threshold pijp_{ij}2 independent of pijp_{ij}3, shows pijp_{ij}4 and pijp_{ij}5 if pijp_{ij}6, and then applies the conditional validity inequality to bound the expected number of false rejections (Wang et al., 2016).

The power advantage over the direct Bonferroni approach comes from replacing the global multiplicity burden pijp_{ij}7 by the filtered set size. Direct Bonferroni for partial conjunction tests rejects pijp_{ij}8 if pijp_{ij}9, which can be extremely conservative because the partial conjunction null is composite. In the pp0 case, if

pp1

with pp2, then applying Bonferroni to pp3 gives

pp4

For large pp5 and sparse signals, this bound is often far below pp6, indicating extreme conservativeness. By contrast, the AdaFilter heuristic for pp7 yields

pp8

which can be much larger than pp9 when P(1)jP(K)jP_{(1)j} \le \cdots \le P_{(K)j}0 (Wang et al., 2016).

AdaFilter-Bon does not satisfy complete monotonicity: reducing some base P(1)jP(K)jP_{(1)j} \le \cdots \le P_{(K)j}1-values can in rare cases reduce the overall threshold and hurt other hypotheses. It does satisfy partial monotonicity: decreasing any of the P(1)jP(K)jP_{(1)j} \le \cdots \le P_{(K)j}2 P(1)jP(K)jP_{(1)j} \le \cdots \le P_{(K)j}3-values for a given hypothesis P(1)jP(K)jP_{(1)j} \le \cdots \le P_{(K)j}4 cannot turn a rejection of P(1)jP(K)jP_{(1)j} \le \cdots \le P_{(K)j}5 into acceptance. The latter is the monotonicity property highlighted as the natural interpretability requirement per hypothesis (Wang et al., 2016).

A sequential extension is also given. After an initial run, remove rejected hypotheses and re-apply AdaFilter-Bon to the remaining ones, iterating until no further rejections. This “Sequential AdaFilter Bonferroni” controls FWER at P(1)jP(K)jP_{(1)j} \le \cdots \le P_{(K)j}6 and increases power (Wang et al., 2016).

5. Practical variants, computation, and empirical behavior

The method accommodates missing P(1)jP(K)jP_{(1)j} \le \cdots \le P_{(K)j}7-values or varying numbers of studies per hypothesis. If hypothesis P(1)jP(K)jP_{(1)j} \le \cdots \le P_{(K)j}8 is measured in only P(1)jP(K)jP_{(1)j} \le \cdots \le P_{(K)j}9 studies and the target replicability level is H0jr/K:fewer than r of the K base hypotheses are non-null.H_{0j}^{r/K}:\quad \text{fewer than } r \text{ of the } K \text{ base hypotheses are non-null}.0, the definitions become

H0jr/K:fewer than r of the K base hypotheses are non-null.H_{0j}^{r/K}:\quad \text{fewer than } r \text{ of the } K \text{ base hypotheses are non-null}.1

The same paper also describes sign-replicability handling by testing one-sided partial conjunction nulls for “H0jr/K:fewer than r of the K base hypotheses are non-null.H_{0j}^{r/K}:\quad \text{fewer than } r \text{ of the } K \text{ base hypotheses are non-null}.2 positives” and “H0jr/K:fewer than r of the K base hypotheses are non-null.H_{0j}^{r/K}:\quad \text{fewer than } r \text{ of the } K \text{ base hypotheses are non-null}.3 negatives” separately at H0jr/K:fewer than r of the K base hypotheses are non-null.H_{0j}^{r/K}:\quad \text{fewer than } r \text{ of the } K \text{ base hypotheses are non-null}.4 each, then taking the union of rejections (Wang et al., 2016).

Computationally, sorting H0jr/K:fewer than r of the K base hypotheses are non-null.H_{0j}^{r/K}:\quad \text{fewer than } r \text{ of the } K \text{ base hypotheses are non-null}.5 H0jr/K:fewer than r of the K base hypotheses are non-null.H_{0j}^{r/K}:\quad \text{fewer than } r \text{ of the } K \text{ base hypotheses are non-null}.6-values for each of H0jr/K:fewer than r of the K base hypotheses are non-null.H_{0j}^{r/K}:\quad \text{fewer than } r \text{ of the } K \text{ base hypotheses are non-null}.7 hypotheses costs H0jr/K:fewer than r of the K base hypotheses are non-null.H_{0j}^{r/K}:\quad \text{fewer than } r \text{ of the } K \text{ base hypotheses are non-null}.8, and the adjusted-H0jr/K:fewer than r of the K base hypotheses are non-null.H_{0j}^{r/K}:\quad \text{fewer than } r \text{ of the } K \text{ base hypotheses are non-null}.9-value implementation requires an additional pass over kk00 and kk01. An R package implementing AdaFilter is available at https://github.com/jingshuw/adaFilter (Wang et al., 2016).

The simulation regime reported for AdaFilter-Bon used kk02, kk03, and sparse signals. At PFER target kk04, direct Bonferroni on partial conjunction kk05-values had PFER kk06–kk07 with recall kk08–kk09 for kk10, and PFER kk11–kk12 with recall kk13–kk14 for kk15. AdaFilter-Bon yielded kk16: PFER kk17–kk18, recall kk19–kk20; and kk21: PFER kk22–kk23, recall kk24–kk25 (Wang et al., 2016).

At FWER target kk26, direct Bonferroni had FWER kk27 with recall kk28–kk29, original AdaFilter-Bon had FWER kk30–kk31 with recall kk32–kk33, and Sequential AdaFilter-Bon had FWER kk34–kk35 with recall kk36–kk37 (Wang et al., 2016).

Although the real applications in the original paper primarily emphasized AdaFilter-BH rather than AdaFilter-Bon, they were presented as corroborating the core AdaFilter advantage across microarray, single-cell RNA sequencing, and metabolomics GWAS settings (Wang et al., 2016).

6. Conservative filtering, kk38-FWER extensions, and terminological ambiguities

A later development analyzed AdaFilter-Bon through the lens of kk39-family-wise error rate. In that formulation, for replication level kk40 across kk41 studies and kk42 features, the generalized threshold is

kk43

and the decision rule rejects kk44 if kk45 (Tran, 21 Aug 2025).

That work also formalized a source of conservativeness induced by filtering. If

kk46

then AdaFilter-Bon satisfies the finite-sample upper bound

kk47

This shows that effective filtering can make the realized error level substantially smaller than kk48, with an accompanying power loss (Tran, 21 Aug 2025).

To mitigate this effect, AdaFilter-AdaBon introduces the post-filter null proportion estimator

kk49

and replaces the rejection threshold by

kk50

Under weak dependence and the stated limit assumptions, the paper proves asymptotic kk51-FWER control and reports higher power than the original AdaFilter-Bon in simulations, with kk52 described as a good default in experiments (Tran, 21 Aug 2025).

A separate source of ambiguity is terminological rather than statistical. In recent language-model work, a similarly spelled “AdaFilter-BoN” refers to adaptive filtering around Best-of-kk53 inference or to consensus filtering for Process Reward Model training, not to partial conjunction multiple testing. “ROC-n-reroll: How verifier imperfection affects test-time scaling” uses “AdaFilter-BoN” for adaptive filtering Best-of-kk54 under an imperfect verifier, with ROC-aware selection of kk55 and thresholds kk56 and switching or combining BoN with rejection sampling (Dorner et al., 16 Jul 2025). “The Lessons of Developing Process Reward Models in Mathematical Reasoning” describes a consensus filtering mechanism that integrates MC estimation with LLM-as-a-judge and states that this is exactly the goal of an “AdaFilter-BoN”-style approach (Zhang et al., 13 Jan 2025). These usages share the idea of adaptive filtering, but they are distinct from AdaFilter-Bon as a Bonferroni-based procedure for partial conjunction testing.

In its original and most specific sense, AdaFilter-Bon denotes a replicability-testing method that combines order-statistic partial conjunction kk57-values with adaptive filtering to control FWER or PFER while substantially reducing the conservativeness of direct Bonferroni correction. Its later refinements chiefly address the residual conservativeness introduced by the filtering step itself (Wang et al., 2016).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to AdaFilter-Bon.