Decision Sciences · MapleScholar Plus

Bridging the P-Value Divide: Extracting Objective Bayesian Evidence from Frequentist Tests

The scientific replication crisis pitted frequentist p-values against Bayesian Bayes factors in bitter methodological warfare; mathematical calibration formulas transform standard p-values into rigorous lower bounds on Bayesian evidence.

Author
Frederik Aust et al.
Published
2026
Journal
arXiv (Cornell University)
Last updated
September 2026
Bridging the P-Value Divide: Extracting Objective Bayesian Evidence from Frequentist Tests

For a century, scientific discovery has relied on frequentist null-hypothesis significance testing, where a p-value below 0.05 is treated as proof of a genuine empirical effect.

Bayesian statisticians fiercely criticized this threshold, demonstrating that a p-value of 0.05 often corresponds to surprisingly weak evidence against the null hypothesis, directly fueling the global scientific replication crisis.

This statistical methodology study establishes rigorous mathematical calibrations that convert standard frequentist p-values into sharp lower bounds on Bayes factors across wide classes of prior distributions. The calibration proves that a p-value of 0.05 represents a maximum odds ratio of only 3-to-1 in favor of the hypothesis.

Translating p-values directly into Bayesian evidence resolves a century-old philosophical feud, equipping medical, psychological, and biological researchers with a unified, transparent standard of empirical proof.

Reference

Aust, F., Pawel, S., & Wagenmakers, E.-J. (2026). Extracting Bayesian Evidence from Frequentist p-Values (Version 1). arXiv.

Title

Extracting Bayesian Evidence from Frequentist p-Values

Abstract

The pp-value and the Bayes factor are measures of evidence that are often considered to be philosophically and mathematically incompatible: The pp-value quantifies conflict between data and H0H_0 ("surprise"), whereas the Bayes factor quantifies the relative predictive accuracy of H0H_0 versus H1H_1 ("evidence"). We revisit Jeffreys's Approximate Bayes factor (JAB) -- a simple, largely overlooked approximation dating back to the 1930s -- which connects these two paradigms for objective hypothesis testing of the existence of an effect. Under a unit-information prior the approximation requires only the pp-value and the effective sample size neffn_\text{eff}. We clarify the core assumptions and boundary conditions for the application of JAB and show across 704 published tt-tests and 39 comparisons of proportions that JAB approximates objective Bayes factors remarkably well. The connection between pp-values and JAB has a practical implication: The evidence implied by a pp-value depends strongly on neffn_\text{eff}. Conventional verbal labels for pp-values (e.g., "strong surprise" for .001 .10canamounttomoderateorevenstrongevidencefor can amount to moderate or even strong evidence for H_0.JABoffersacheap,sample−size−sensitivesupplementto. JAB offers a cheap, sample-size-sensitive supplement to p$-values, computable from routinely reported statistics, that remains valid even under optional stopping.

Cited 0 times · View on doi.org

Continue

Continue Exploring

Ask this paper your own questions, or keep browsing the verified research catalogue.