The scientific replication crisis pitted frequentist p-values against Bayesian Bayes factors in bitter methodological warfare; mathematical calibration formulas transform standard p-values into rigorous lower bounds on Bayesian evidence.

For a century, scientific discovery has relied on frequentist null-hypothesis significance testing, where a p-value below 0.05 is treated as proof of a genuine empirical effect.
Bayesian statisticians fiercely criticized this threshold, demonstrating that a p-value of 0.05 often corresponds to surprisingly weak evidence against the null hypothesis, directly fueling the global scientific replication crisis.
This statistical methodology study establishes rigorous mathematical calibrations that convert standard frequentist p-values into sharp lower bounds on Bayes factors across wide classes of prior distributions. The calibration proves that a p-value of 0.05 represents a maximum odds ratio of only 3-to-1 in favor of the hypothesis.
Translating p-values directly into Bayesian evidence resolves a century-old philosophical feud, equipping medical, psychological, and biological researchers with a unified, transparent standard of empirical proof.
Extracting Bayesian Evidence from Frequentist p-Values
The -value and the Bayes factor are measures of evidence that are often considered to be philosophically and mathematically incompatible: The -value quantifies conflict between data and ("surprise"), whereas the Bayes factor quantifies the relative predictive accuracy of versus ("evidence"). We revisit Jeffreys's Approximate Bayes factor (JAB) -- a simple, largely overlooked approximation dating back to the 1930s -- which connects these two paradigms for objective hypothesis testing of the existence of an effect. Under a unit-information prior the approximation requires only the -value and the effective sample size . We clarify the core assumptions and boundary conditions for the application of JAB and show across 704 published -tests and 39 comparisons of proportions that JAB approximates objective Bayes factors remarkably well. The connection between -values and JAB has a practical implication: The evidence implied by a -value depends strongly on . Conventional verbal labels for -values (e.g., "strong surprise" for .001 .10H_0p$-values, computable from routinely reported statistics, that remains valid even under optional stopping.
Ask this paper your own questions, or keep browsing the verified research catalogue.