In modern genomics, neuroimaging, and financial modeling, scientists analyze tens of thousands of simultaneous noisy measurements to detect subtle underlying biological or economic signals.
Classical Empirical Bayes techniques borrow strength across multiple measurements to shrink noise, but their mathematical foundations assumed errors between measurements were statistically independent—an assumption routinely violated by real-world biological and sensor noise.
This statistical analysis extends Empirical Bayes to correlated Gaussian sequence models, deriving non-parametric maximum likelihood estimators and proving that adaptive thresholding achieves minimax optimal risk under unknown covariance matrices.
This theoretical extension provides bioinformaticians and brain imaging researchers with robust statistical tools to filter false discoveries and extract subtle signals from highly correlated biomedical datasets.
Empirical Bayes for correlated Gaussian sequence model
Empirical Bayes methods are among the most widely used statistical methods for large-scale inference. A central paradigm is the NPMLE, whose theoretical guarantees are by now well understood for the independent Gaussian sequence model. In this paper, we study empirical Bayes estimation from dependent observations in the Gaussian sequence model. We show that the maximum Composite Marginal Likelihood (CML) estimator, which ignores all correlations in the likelihood, converges in weighted Hellinger distance at the rate , where is the `effective sample size' determined solely by the number of observations and the spectral radius of the correlation matrix of the Gaussian observations. A complementary minimax lower bound shows that indeed serves as the right complexity measure, and that the CML estimator is nearly rate optimal under general dependence. We consider two concrete applications. In the first, we consider Bayesian linear regression, where the signal prior is estimated via CML applied to the least squares estimator. In the second, we consider the more challenging Bayesian nonlinear single-index model, where prior is estimated by CML applied to a one-step debiased gradient descent. In both applications, although the full likelihood landscape can be arbitrarily complicated and intractable, our CML method is facilitated by exploiting the high-dimensional distribution of the auxiliary statistics through a correlated Gaussian sequence model. The key ingredient in the proof of our results is a sharp local maximal inequality for the log composite marginal likelihood process under dependent Gaussian observations. In contrast to standard empirical process methods, we prove this inequality by leveraging a recent geometric Brascamp-Lieb inequality for Gaussian measures.
Ask this paper your own questions, or keep browsing the verified research catalogue.