Computer Science · MapleScholar Plus

Decoding Correlated Noise: Empirical Bayes for High-Dimensional Gaussian Sequences

Empirical Bayes methods traditionally assumed observational noise errors were completely independent and uncorrelated; this statistical breakthrough establishes minimax optimal estimation for high-dimensional Gaussian sequences under general correlation structures.

Author
Qiyang Han et al.
Published
2026
Journal
arXiv (Cornell University)
Last updated
September 2026
Decoding Correlated Noise: Empirical Bayes for High-Dimensional Gaussian Sequences

In modern genomics, neuroimaging, and financial modeling, scientists analyze tens of thousands of simultaneous noisy measurements to detect subtle underlying biological or economic signals.

Classical Empirical Bayes techniques borrow strength across multiple measurements to shrink noise, but their mathematical foundations assumed errors between measurements were statistically independent—an assumption routinely violated by real-world biological and sensor noise.

This statistical analysis extends Empirical Bayes to correlated Gaussian sequence models, deriving non-parametric maximum likelihood estimators and proving that adaptive thresholding achieves minimax optimal risk under unknown covariance matrices.

This theoretical extension provides bioinformaticians and brain imaging researchers with robust statistical tools to filter false discoveries and extract subtle signals from highly correlated biomedical datasets.

Reference

Han, Q., & Zhang, C.-H. (2026). Empirical Bayes for correlated Gaussian sequence model (Version 1). arXiv.

Title

Empirical Bayes for correlated Gaussian sequence model

Abstract

Empirical Bayes methods are among the most widely used statistical methods for large-scale inference. A central paradigm is the NPMLE, whose theoretical guarantees are by now well understood for the independent Gaussian sequence model. In this paper, we study empirical Bayes estimation from dependent observations in the Gaussian sequence model. We show that the maximum Composite Marginal Likelihood (CML) estimator, which ignores all correlations in the likelihood, converges in weighted Hellinger distance at the rate n∗−1/2n_*^{-1/2}, where n∗=n/κ0n_*=n/κ_0 is the `effective sample size' determined solely by the number of observations nn and the spectral radius κ0κ_0 of the correlation matrix of the Gaussian observations. A complementary minimax lower bound shows that n∗n_* indeed serves as the right complexity measure, and that the CML estimator is nearly rate optimal under general dependence. We consider two concrete applications. In the first, we consider Bayesian linear regression, where the signal prior is estimated via CML applied to the least squares estimator. In the second, we consider the more challenging Bayesian nonlinear single-index model, where prior is estimated by CML applied to a one-step debiased gradient descent. In both applications, although the full likelihood landscape can be arbitrarily complicated and intractable, our CML method is facilitated by exploiting the high-dimensional distribution of the auxiliary statistics through a correlated Gaussian sequence model. The key ingredient in the proof of our results is a sharp local maximal inequality for the log composite marginal likelihood process under dependent Gaussian observations. In contrast to standard empirical process methods, we prove this inequality by leveraging a recent geometric Brascamp-Lieb inequality for Gaussian measures.

Cited 0 times · View on doi.org

Continue

Continue Exploring

Ask this paper your own questions, or keep browsing the verified research catalogue.