AdaFilter-AdaBon: Adaptive Multiple Testing
- The paper refines the AdaFilter-Bon framework by introducing a post-filter null proportion estimator to mitigate conservativeness in detecting replicated signals.
- It employs filtering of partial conjunction p-values to identify features replicated across studies while controlling the generalized k-FWER.
- The method demonstrates enhanced power with asymptotic error guarantees and scalable O(m) computational complexity for high-dimensional analyses.
Searching arXiv for the cited papers and closely related work to ground the article. AdaFilter-AdaBon is an adaptive multiple-testing procedure for detecting replicated signals across several studies while controlling a generalized family-wise error rate, the -FWER. It builds directly on the AdaFilter-Bon method of Wang et al. (2022), and improves its power by estimating the proportion of true nulls after filtering (Tran, 21 Aug 2025). In the literature, the name “AdaFilter” is also used for an unrelated adaptive fine-tuning method in deep transfer learning; that 2019 convolutional-filter method is distinct from the statistical AdaFilter line introduced for partial conjunction testing and should not be conflated with AdaFilter-AdaBon (Guo et al., 2019).
1. Historical lineage and nomenclature
AdaFilter-AdaBon belongs to a line of methods for partial conjunction (PC) hypotheses, where the scientific objective is to identify signals that replicate across multiple studies rather than signals that are merely significant in an aggregate meta-analysis. The earlier AdaFilter framework introduced adaptive filtering procedures for PC hypotheses and developed two principal variants: AdaFilter Bonferroni, often referred to informally as “AdaBon,” for FWER/PFER control, and AdaFilter BH for FDR control (Wang et al., 2016). AdaFilter-AdaBon is a later adaptive refinement of the Bonferroni branch: it retains the same filtering logic as AdaFilter-Bon, but modifies the threshold by incorporating a post-filter estimate of the null proportion (Tran, 21 Aug 2025).
The statistical use of “AdaFilter” is unrelated to the computer-vision method “AdaFilter: Adaptive Filter Fine-tuning for Deep Transfer Learning,” which is an adaptive filter-level fine-tuning framework for deep transfer learning and operates on duplicated convolutional filters and recurrent gating in ResNet-50-style architectures (Guo et al., 2019). That naming overlap is purely terminological. In the multiple-testing literature, “AdaBon” denotes the Bonferroni-style adaptive filtering procedure derived from the original AdaFilter framework (Wang et al., 2016), whereas “AdaFilter-AdaBon” denotes the later procedure that estimates the post-filter null proportion to mitigate conservativeness (Tran, 21 Aug 2025).
A common misconception is therefore to read “AdaFilter-AdaBon” as a hybrid of deep-learning fine-tuning and statistical error control. The record in the cited papers supports the opposite conclusion: AdaFilter-AdaBon is a multiple-testing procedure for replicability analysis, while the 2019 CNN AdaFilter is a transfer-learning algorithm with no direct methodological connection (Tran, 21 Aug 2025).
2. Statistical setting: partial conjunction, replicability, and -FWER
The method considers a meta-analysis of comparable studies, each measuring the same features. For feature in study , there is a null hypothesis and a corresponding valid -value , meaning
Across studies, the vectors 0 are assumed independent, while within each study there can be dependence across features (Tran, 21 Aug 2025).
Replicability is encoded by a number 1. A feature is “replicated at level 2” if at least 3 of its study-specific nulls are false. The corresponding PC null is
4
with the alternative that at least 5 of the 6 component nulls are false (Tran, 21 Aug 2025). Rejecting 7 therefore means declaring feature 8 replicated in at least 9 studies.
A key structural property is nesting: 0 This nesting is central to the filtering logic, because failure to replicate at level 1 precludes replication at level 2 (Tran, 21 Aug 2025).
For testing, the method uses the Bonferroni PC 3-value. If 4 are the order statistics of the study-specific 5-values, then
6
This is the choice used inside AdaFilter-Bon and AdaFilter-AdaBon (Tran, 21 Aug 2025). The focus is not the ordinary FWER alone, but the generalized 7-family-wise error rate,
8
which allows up to 9 false rejections. When 0, this coincides with FWER (Tran, 21 Aug 2025).
The motivation is the classical multiplicity burden of high-dimensional replicability analysis. Even when one has valid PC 1-values, testing 2 hypotheses with FWER or 3-FWER control leads to very stringent per-hypothesis thresholds. This motivates procedures that reduce the multiplicity burden without sacrificing error control, and filtering is the specific device used here (Tran, 21 Aug 2025).
3. Construction of AdaFilter-AdaBon
AdaFilter-AdaBon keeps the filtering structure of AdaFilter-Bon but adapts the rejection threshold using an estimate of the post-filter null proportion (Tran, 21 Aug 2025). For a fixed replicability level 4, it defines two quantities for each feature 5: 6 the testing PC 7-value, and
8
the filtering 9-value, obtained from the PC null at level 0 (Tran, 21 Aug 2025). Because 1, one has 2.
In AdaFilter-Bon, the rejection threshold is
3
so the multiplicity factor is the number retained by the filter rather than the full 4 (Tran, 21 Aug 2025). The later AdaFilter-AdaBon procedure modifies this by estimating the proportion of true nulls among the retained features. Fix a tuning parameter 5. The estimator is
6
Its numerator counts retained features whose PC 7-value is at least 8, and its denominator normalizes by the retained set size and the factor 9 (Tran, 21 Aug 2025).
The conceptual AdaFilter-AdaBon threshold is
0
Plugging in the estimator yields the operational definition: 1 The rejection rule is then simply: reject 2 if 3 (Tran, 21 Aug 2025).
The paper also gives an implementable finite grid
4
and defines
5
A key theorem states that the rejection sets coincide: 6 The paper states that one can compute 7 in 8 time by evaluating the left-hand side at each candidate in 9 (Tran, 21 Aug 2025).
This construction preserves the original AdaFilter idea—screen with 0, test with 1—but replaces the fixed post-filter multiplicity correction of AdaFilter-Bon with a multiplicity correction scaled by an estimated post-filter null proportion. This suggests a direct mechanism for reducing conservativeness when filtering preferentially retains alternatives.
4. Theoretical guarantees and regularity conditions
The main theoretical result is asymptotic 2-FWER control under weak dependence assumptions (Tran, 21 Aug 2025). The paper defines 3 as the number of true PC nulls and 4 as the number of alternatives, and assumes almost-sure convergence of the empirical distributions of 5 and 6 under both the null and the alternative. It also assumes
7
These are presented as typical empirical-process convergence assumptions encompassing many weak-dependence structures, including finite blocks and mixing processes (Tran, 21 Aug 2025).
A central inequality is
8
which reflects a conditional validity lemma: under the PC null and independence of the study 9-values,
0
This is the asymptotic analogue of the conditional validity phenomenon that already underpinned the original AdaFilter theory (Wang et al., 2016).
Let
1
be the rejection set and
2
the number of false discoveries. The main theorem states that if Assumption 1 holds, if 3 for some 4, and if there exists 5 such that
6
with probability 7, then
8
The paper’s interpretation is that the estimator 9 is constructed so that, under the null, the proportion of retained features with 0 is at least 1 in expectation. Intuitively, this makes 2 a conservative estimate of 3 in large samples, so using it in the threshold tightens the constraint when necessary rather than weakening it unsafely (Tran, 21 Aug 2025).
The guarantees are explicitly asymptotic. The paper states that no explicit finite-sample guarantees are provided, although simulations show good empirical control across a range of settings and correlations (Tran, 21 Aug 2025). This is an important qualification: AdaFilter-AdaBon is theoretically rigorous, but its formal guarantee is not a finite-sample exact 4-FWER theorem in the style of the original AdaFilter Bonferroni result under full independence (Wang et al., 2016).
5. Power improvement, simulations, and computational profile
The motivation for AdaFilter-AdaBon is the conservativeness of AdaFilter-Bon. The paper gives the bound
5
for AdaFilter-Bon under independent and valid 6 (Tran, 21 Aug 2025). Because filtering preferentially retains hypotheses with signal, the post-filter null proportion 7 can be much smaller than 8, so AdaFilter-Bon may control at a level far below the nominal 9. AdaFilter-AdaBon is designed to mitigate exactly this conservativeness.
The paper’s heuristic comparison is explicit. Under AdaFilter-Bon, the effective threshold is roughly
0
where 1 is the number retained by the filter. Under AdaFilter-AdaBon, the effective threshold is roughly
2
If 3 correctly estimates the post-filter null proportion, the expected number of false rejections becomes approximately 4, and the threshold is larger by approximately a factor of 5 (Tran, 21 Aug 2025). This suggests that the method is specifically advantageous when filtering removes a substantial fraction of true nulls.
The reported simulations use 6 features and 7 studies, with blockwise equicorrelated Gaussian noise within each study, 8, signal density 9, replicability levels 00, and target FWER (01) at 02 (Tran, 21 Aug 2025). Compared methods include AdaFilter-AdaBon (03), AdaFilter-Bon, standard Bonferroni and Hochberg applied to Fisher PC 04-values, and adaptive Bonferroni and adaptive Hochberg applied to Fisher PC 05-values.
The stated findings are that AdaFilter-AdaBon maintains FWER below 06 across all settings; its FWER is systematically higher than AdaFilter-Bon but still comfortably below target; and it yields notably higher TPR than AdaFilter-Bon in most settings, especially for 07 and moderate-to-high signal density 08 (Tran, 21 Aug 2025). For very stringent replication (09) and extremely sparse signals (10), AdaFilter-Bon matches AdaFilter-AdaBon’s power. Additional simulations for 11 and 12 show that all methods are very conservative, while AdaFilter-AdaBon is the least conservative among them and has substantially higher power, especially when 13 (Tran, 21 Aug 2025).
On implementation, the paper states that sorting the study 14-values per feature is 15 per feature, but because 16 is small, the overall complexity is 17 in practice; constructing 18 is 19; and evaluating the adaptive statistic can be made linear by cumulative counts (Tran, 21 Aug 2025). The method therefore scales linearly in 20 up to large numbers of features.
6. Relation to AdaFilter, AdaBon, and adjacent procedures
The original AdaFilter formulation introduced the filtering and selection statistics
21
and used them to define the AdaFilter Bonferroni threshold
22
for FWER/PFER control (Wang et al., 2016). In that sense, AdaFilter-AdaBon is not a new filtering architecture but a modification of the calibration step: the term 23 remains, but it is multiplied by an adaptive estimate of the post-filter null proportion (Tran, 21 Aug 2025).
This relation is best understood in three layers. First, AdaFilter is the general framework for adaptive filtering in PC testing (Wang et al., 2016). Second, AdaBon or AdaFilter Bonferroni is the Bonferroni-style instantiation of that framework (Wang et al., 2016). Third, AdaFilter-AdaBon is the later procedure that augments AdaFilter-Bon with post-filter null-proportion estimation to improve power while preserving asymptotic 24-FWER control (Tran, 21 Aug 2025).
The paper also positions AdaFilter-AdaBon relative to standard Bonferroni, Holm, and Hochberg procedures for PC hypotheses. Those methods treat the PC 25-values like ordinary single-study 26-values and penalize all 27 hypotheses equally; they do not use filtering and do not use null-proportion estimation; and they are therefore described as severely conservative in high-dimensional replicability analysis (Tran, 21 Aug 2025). In the broader replicability literature, two-stage selection-and-testing procedures and FDR-based partial conjunction methods address related questions, but they do not estimate post-filter null proportions in the same way (Tran, 21 Aug 2025).
The paper further notes that AdaFilter-AdaBon is defined using Bonferroni PC 28-values. If one prefers more powerful combining functions such as Fisher, one cannot directly plug them into this specific algorithm; adapting the method to other combining functions would require new theory (Tran, 21 Aug 2025). This is a substantive methodological boundary rather than a mere implementation detail.
Finally, the 2025 paper outlines extensions beyond 29-FWER. It states that AdaFilter-AdaBon can be augmented to control the false exceedance rate (FDX) and the false discovery rate (FDR) asymptotically through a second threshold 30 and a Genovese–Wasserman-style bound (Tran, 21 Aug 2025). A plausible implication is that the adaptive-filtering-plus-post-filter-estimation principle is broader than the specific 31-FWER instantiation, although the developed theory in the paper is centered on replicability analysis with Bonferroni PC 32-values.
7. Interpretation, scope, and limitations
A rejection by AdaFilter-AdaBon means that the corresponding feature is declared replicated in at least 33 studies, with the global error metric controlled at the specified 34-FWER level in the asymptotic sense established by the paper (Tran, 21 Aug 2025). This is a stronger statement than ordinary meta-analytic significance, because the null explicitly concerns the number of studies in which the effect is non-null.
The method is expected to be most beneficial when there are many features, when a nontrivial fraction of features are truly replicated at level 35, when the replication requirement 36 is modest, and when dependence across features is moderate rather than pathological (Tran, 21 Aug 2025). This follows the paper’s discussion that filtering is especially informative when 37 replication serves as a meaningful screen for 38-level replication.
Several limitations are explicit. The asymptotic theory allows certain forms of weak dependence across features, but very strong or complex dependence structures may violate the assumptions (Tran, 21 Aug 2025). The guarantees are asymptotic as 39, so for small 40 the estimator 41 may be noisy and the adaptive threshold may have less predictable behavior (Tran, 21 Aug 2025). The formal theory is developed for Bonferroni PC 42-values, and the asymptotic 43-FWER theorem assumes that 44 grows linearly with 45, even though simulations examine fixed small 46 such as 47, 48, and 49 (Tran, 21 Aug 2025).
These limitations do not negate the method’s contribution; they define its scope. In the statistical literature, AdaFilter-AdaBon is best understood as an adaptive refinement of AdaFilter-Bon for replicated-signal detection under partial conjunction testing, with asymptotic 50-FWER control and empirically higher power than the original AdaFilter-Bon (Tran, 21 Aug 2025). In contrast, the identically named but unrelated deep-learning AdaFilter remains a transfer-learning procedure on convolutional networks and has no role in the statistical construction of AdaFilter-AdaBon (Guo et al., 2019).