Multiple Testing, Multilevel Solutions: Revisiting The Multiplicity Problem With Frequentist And Bayesian Hierarchical Modeling

Loading...
Thumbnail Image

Authors

Alter, Udi

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

The longstanding multiplicity problem—the inflation of false positives that arises from conducting multiple significance tests—remains a source of confusion and disagreement in psychology and many other sciences. Traditional solutions, such as familywise error rate (FWE; e.g., Bonferroni) and false discovery rate (FDR; e.g., Benjamini-Hochberg) corrections, impose control externally by adjusting p values and confidence intervals. These post-hoc adjustments can reduce false positives but often at the expense of statistical power, making it harder to detect true effects when many hypotheses are tested simultaneously. Hierarchical modeling incorporates multiplicity as part of the model itself and allows information to be shared appropriately across tests. As a result, the inferences drawn about the effect sizes, uncertainty intervals, and statistical significance are often more justifiable. The first simulation study evaluated frequentist multilevel modeling (MLM) as an alternative to post-hoc adjustment procedures. Results provided initial evidence that partial pooling and shrinkage can help address the multiplicity problem, particularly when cluster sizes and effect magnitudes are large. Although MLM is not a universal solution, it offers a substitute for mechanical corrections, but only under certain conditions. The second simulation study extended the investigation into Bayesian multilevel modeling (BMLM). Findings were mixed: when the joint null hypothesis was true (i.e., all effects were null), most BMLM prior variants maintained very low false-positive rates—well below the nominal Type I error rate and other correction methods—while retaining high power. Yet in partial-null settings, where true and null effects coexisted, strong shrinkage and narrow credible intervals raised familywise error above that of MLM, FWE, and FDR. Together, the two studies demonstrate that hierarchical models can, under the right conditions, embed multiplicity control within the model itself rather than impose it after the fact. However, their performance depends critically on the structural and design conditions, effect configuration, and modeling assumptions. This work adds empirical evidence to ongoing theoretical discussions of how multilevel and Bayesian approaches can promote more transparent, coherent, and defensible inference.

Description

Keywords

Quantitative psychology and psychometrics, Psychology, Statistics

Citation