Post-selection Inference in Multiverse Analysis (PIMA): an inferential framework based on the sign flipping score test
Abstract: When analyzing data researchers make some decisions that are either arbitrary, based on subjective beliefs about the data generating process, or for which equally justifiable alternative choices could have been made. This wide range of data-analytic choices can be abused, and has been one of the underlying causes of the replication crisis in several fields. Recently, the introduction of multiverse analysis provides researchers with a method to evaluate the stability of the results across reasonable choices that could be made when analyzing data. Multiverse analysis is confined to a descriptive role, lacking a proper and comprehensive inferential procedure. Recently, specification curve analysis adds an inferential procedure to multiverse analysis, but this approach is limited to simple cases related to the linear model, and only allows researchers to infer whether at least one specification rejects the null hypothesis, but not which specifications should be selected. In this paper we present a Post-selection Inference approach to Multiverse Analysis (PIMA) which is a flexible and general inferential approach that accounts for all possible models, i.e., the multiverse of reasonable analyses. The approach allows for a wide range of data specifications (i.e. pre-processing) and any generalized linear model; it allows testing the null hypothesis of a given predictor not being associated with the outcome, by merging information from all reasonable models of multiverse analysis, and provides strong control of the family-wise error rate such that it allows researchers to claim that the null-hypothesis can be rejected for each specification that shows a significant effect. The inferential proposal is based on a conditional resampling procedure. To be continued...
- Agresti, A. (2015). Foundations of Linear and Generalized Linear Models. Wiley, New York.
- Publication bias: a problem in interpreting medical data. Journal of the Royal Statistical Society: Series A (Statistics in Society), 151(3):419–445.
- Benjamini, Y. (2020). Selective Inference: The Silent Killer of Replicability. Harvard Data Science Review, 2(4). https://hdsr.mitpress.mit.edu/pub/l39rpgyc.
- Berger, R. L. (1982). Multiparameter hypothesis testing and acceptance sampling. Technometrics, 24(4):295–300.
- Star wars: The empirics strike back. American Economic Journal: Applied Economics, 8(1):1–32.
- Associations of covid-19 risk perception with vaccine hesitancy over time for italian residents. Social Science & Medicine, 272:113688.
- Inference in generalized linear models with robustness to misspecified variances.
- Increasing the transparency of research papers with explorable multiverse analyses. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, pages 1–15.
- Systematic review of the empirical evidence of study publication bias and outcome reporting bias. PloS one, 3(8):e3081.
- Fanelli, D. (2012). Negative results are disappearing from most disciplines and countries. Scientometrics, 90(3):891–904.
- Finner, H. (1999). Stepwise multiple test procedures and control of directional errors. The Annals of Statistics, 27(1):274–289.
- Finos, L. (2022). jointest: Multivariate testing trought joint sign-flip scores (Hemerik, Goeman and Finos (2020) ¡doi:10.1111/rssb.12369¿). R package version 1.2.0.
- Fisher, R. A. (1925). Statistical Methods for Research Workers. Oliver and Boyd, Edinburgh.
- Flachaire, E. (1999). A better way to bootstrap pairs. Economics Letters, 64(3):257–262.
- Freedman, D. A. (1981). Bootstrapping regression models. The Annals of Statistics, 9(6):1218–1228.
- Identifying robust correlates of risk preference: A systematic approach using specification curve analysis. Journal of Personality and Social Psychology, 120(2):538.
- The statistical crisis in science data-dependent analysis–a ”garden of forking paths” – explains why many statistically significant comparisons don’t hold up. American scientist, 102(6):460.
- Exceedance control of the false discovery proportion. J. Am. Statist. Ass., 101(476):1408–1417.
- Only closed testing procedures are admissible for controlling false discovery proportions. Ann. Statist., 49(2):1218–1238.
- Multiple testing for exploratory research. Statist. Sci., 26(4):584–597.
- Greenwald, A. G. (1975). Consequences of prejudice against the null hypothesis. Psychological bulletin, 82(1):1.
- Harder, J. A. (2020). The multiverse of methods: Extending the multiverse analysis to address data-collection decisions. Perspectives on Psychological Science, 15(5):1158–1177.
- Exact testing with random permutations. TEST, 27:811–825.
- Robust testing in generalized linear models by sign flipping score contributions. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 82(3):841–864.
- Examining the robustness of observational associations to model, measurement and sampling uncertainty with the vibration of effects framework. International Journal of Epidemiology, 50(1):266–278.
- Liptak, T. (1958). On the combination of independent tests. Magyar Tud. Akad. Mat. Kutató Int. Közl., 3:1971–1977.
- Boba: Authoring and visualizing multiverse analyses. IEEE Transactions on Visualization and Computer Graphics, 27(2):1753–1763.
- On closed testing procedures with special reference to ordered analysis of variance. Biometrika, 63(3):655–660.
- Advancing our understanding of cognitive development and motor vehicle crash risk: a multiverse representation analysis. Cortex, 138:90–100.
- Tuning into the real effect of smartphone use on parenting: a multiverse analysis. Journal of Child Psychology and Psychiatry, 61(8):855–865.
- A method to increase the credibility of published results. Social Psychology, 45(3):137–141.
- Open Science Collaboration (2015). Estimating the reproducibility of psychological science, volume 349. American Association for the Advancement of Science.
- Pesarin, F. (2001). Multivariate Permutation Tests: with Applications in Biostatistics. Wiley, New York.
- R Core Team (2021). R: A Language and Environment for Statistical Computing. R Foundation for Statistical Computing, Vienna, Austria.
- Permutation tests using arbitrary permutation distributions. Sankhya A, pages 1–22.
- Assessing the robustness of mediation analysis results using multiverse analysis. Prevention Science, pages 1–11.
- Shaffer, J. P. (1980). Control of directional errors with stagewise multiple test procedures. The Annals of Statistics, 8(6):1342–1347.
- False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological science, 22(11):1359–1366.
- Specification curve analysis. Nature Human Behaviour, 4(11):1208–1214.
- Increasing transparency through a multiverse analysis. Perspectives on Psychological Science, 11(5):702–712.
- Sterling, T. D. (1959). Publication decisions and their possible effects on inferences drawn from tests of significance—or vice versa. Journal of the American Statistical Association, 54(285):30–34.
- van der Vaart, A. W. (1998). Asymptotic Statistics. Cambridge University Press, Cambridge.
- Permutation-based true discovery guarantee by sum tests. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 85(3).
- Resampling-based multisplit inference for high-dimensional regression. https://arxiv.org/abs/2205.12563.
- A multiverse analysis of early attempts to replicate memory suppression with the think/no-think task. Memory, 28(7):870–887.
- Resampling-Based Multiple Testing: Examples and Methods for p-Value Adjustment. Wiley, New York.
Paper Prompts
Sign up for free to create and run prompts on this paper.