- The paper introduces PIMAX, a novel framework that extends multiverse analysis to GLMMs using sign-flipping score tests and two-stage summary-statistics.
- It achieves strong family-wise error rate control without needing explicit random-effects specification, ensuring valid inference in clustered data.
- Simulation studies and a SHARE data analysis demonstrate that PIMAX outperforms traditional GLMM methods in type I error control and statistical power.
Post-Selection Inference for Multiverse Analysis in Mixed-Effects Models (PIMAX)
Overview
The proliferation of analytic flexibility in high-dimensional and clustered data motivates robust, transparent, and multiplicity-aware inference procedures. This paper introduces PIMAX, a post-selection inferential framework that extends valid multiverse analysis to generalized linear mixed models (GLMMs), leveraging sign-flipping score tests and two-stage summary-statistics methodology. PIMAX achieves FWER control across a user-specified model multiverse in clustered data scenarios, circumventing the pitfalls of random-effects misspecification and the attendant inflation of type I error rates endemic to conventional GLMM procedures.
Technical Contributions
The methodological innovation of PIMAX is the synthesis of the PIMA framework for post-selection inference in GLMs with the flip2sss procedure for robust cluster-level inference. The composition yields:
- Asymptotically valid inference for the global null and for specification-level hypotheses across a user-defined multiverse.
- Strong FWER control via closed testing and max-type combination functions.
- Elimination of the need for explicit random-effects structure specification, thereby maintaining error control under heteroscedasticity, unbalanced designs, and within-cluster dependence.
- Empirical validation against GLMM-based approaches, demonstrating both rigorous type I error behavior and high statistical power.
Clustered Data and Specification Multiverses
The inferential target in clustered data frequently involves uncertainty both in the fixed-effects design (covariate inclusion, transformations) and in the random-effects structure (e.g., random slopes selection, covariance parameterization). PIMAX is agnostic to the precise random-effects distribution and instead reduces the data to cluster-level summaries, effecting a semiparametric transfer of the inferential burden to nuisance-robust quantities.
Formally, for K models in the multiverse, cluster-specific summary statistics are generated in a first-stage GLM, then analyzed in a second-stage working model. This decouples random-effects uncertainty from fixed-effects selection, ensuring that classical post-selection tools from GLM contexts can be imported to clustered-data settings.
Sign-Flipping Score-Based Testing
The core of the methodology is the sign-flipping score test, implemented at the cluster summary level, providing valid inference under minimal moment and independence assumptions at the cluster level. The sign-flipping mechanism builds a null reference by randomly switching the signs of cluster-specific scores, yielding an estimated null distribution for the standardized aggregate statistic. This resampling mechanism does not require normality or homoskedasticity.
PIMAX extends the sign-flipping paradigm to multiverse settings by applying common sign-flipping patterns across all candidate specifications. This preserves the true dependence structure among model-specific test statistics—a necessity for valid multiplicity corrections.
Framework for Multiverse Inference
For the multiverse M, PIMAX provides three inferential outputs:
- Global test: A combined statistic (typically using ψ=max or ψ=mean) yields a global p-value for the intersection null of all specifications.
- Adjusted specification-level inference: Closed testing delivers multiplicity-adjusted p-values, guaranteeing strong FWER control across the multiverse.
- Simultaneous lower confidence bounds: On the number of non-null effects, via min-max based procedures.
This is operationalized via the application of the same B sign-flip transformations to all specifications, aggregating the resultant scores with predefined, monotonic combination functions.

Figure 1: Empirical type I error of PIMAX across simulation settings (within-cluster scenario); the method adheres to nominal level α=0.05.

Figure 2: Type I error rates in the between-cluster scenario demonstrate that GLMM-based inference is anti-conservative while PIMAX remains controlled.
Simulation Results
Empirical studies with binary outcomes simulated under controlled cluster structures reveal that PIMAX adheres to the nominal type I error rate in both within- and between-cluster settings. In contrast, GLMM-based inference—even with Holm-Bonferroni adjustment—regularly exhibits inflated type I error, particularly with model misspecification or small cluster counts.

Figure 3: Statistical power for detecting nonzero effects in the within-cluster scenario; the mean-combination variant of PIMAX exhibits consistently maximal power.

Figure 4: Power analysis in the between-cluster scenario where GLMM-based inference fails type I error control and is omitted; PIMAX retains power and validity.
For global hypothesis tests, the mean-based combining function for PIMAX yields maximal sensitivity, while max-based combining functions (and closed-testing maxT shortcuts) offer balanced error rates and interpretable multiplicity control.
Application: SHARE Data Analysis
The methodology is applied to the SHARE (Survey of Health, Ageing, and Retirement in Europe) database, with 48 model specifications arising from defensible choices in operationalizing predictors and outcomes. Hypothesis tests on all covariates are performed across multiverse specifications, and the method's conservative post-selection inference restricts significant findings predominantly to robust demographic predictors (age and sex), revealing the instability and model-dependence of socio-economic and health-related associations.

Figure 5: Distributions of estimated chronic morbidity effects versus −log10(p) across multiverse specifications, emphasizing variability and the influence of specification choice on inference.
The analysis substantiates the necessity of multiverse-aware inference: effect estimates and their statistical significance can be highly sensitive to analytic choices; valid inference demands correction for researcher degrees of freedom.
Implications and Future Directions
PIMAX lays the groundwork for robust, multiplicity-aware inference in multiverse settings with clustered data. Theoretically, it integrates resampling-based inference with strong FWER control schemes, overcoming the limitations of both classical mixed-models inference (sensitive to random-effects misspecification) and post-selection methods confined to i.i.d. data. Practically, analysts can produce valid p-values and confidence statements after flexible analytic exploration, without risk of overstating evidence.
Extensions could include adaptation to more complex correlation structures (e.g., crossed random effects, spatial/temporal hierarchies), development of more powerful score aggregation mechanisms for structured multiverses, and further optimization of cluster-level summary estimators, particularly in small-M0 or sparse-cluster regimes.
Conclusion
PIMAX delivers a unified approach for post-selection inference in multiverse analysis with clustered data. By bypassing the need for explicit parametric random-effects modeling, leveraging sign-flipping mechanics, and implementing closed-testing algorithms, it ensures valid FWER control and interpretable power for analysts navigating model uncertainty. Simulation and applied evidence validate its superiority to conventional alternatives under clustered dependence and multiplicity, providing a methodological foundation for principled scientific inference in settings plagued by both model and sampling complexity.