Chat
Psychology · MapleScholar Plus

Disentangling Noisy Labels: Amortized Variational Inference for Partial-Label Learning

Crowdsourced and medical datasets frequently assign multiple conflicting candidate labels to a single instance; amortized variational inference models true latent labels as unobserved variables to achieve robust machine learning.

Author
Fuchs, Tobias et al.
Published
2025
Journal
arXiv (Cornell University)
Last updated
September 2026
Disentangling Noisy Labels: Amortized Variational Inference for Partial-Label Learning

Supervised deep learning depends on vast collections of accurately labeled training data, yet collecting pristine ground-truth annotations across medicine, biology, and computer vision is expensive and prone to human error.

In crowdsourcing and web-scraped corpora, annotators regularly assign multiple plausible candidate labels per sample (partial-label learning), leaving algorithms unable to discern the true target from distracting label noise.

This machine learning study establishes an amortized variational inference framework that models the true class label as a latent variable. By parameterizing the approximate posterior through a neural inference network, the model jointly optimizes label disambiguation and classifier training in an end-to-end Bayesian pipeline.

Amortized partial-label learning unlocks high-accuracy model training on cheap, ambiguous real-world data, drastically reducing annotation costs across clinical pathology, satellite remote sensing, and autonomous vision.

Reference

Fuchs, T., & Klein, N. (2025). Amortized Variational Inference for Partial-Label Learning: A Probabilistic Approach to Label Disambiguation (Version 2). arXiv.

Title

Amortized Variational Inference for Partial-Label Learning: A Probabilistic Approach to Label Disambiguation

Abstract

Real-world data is frequently noisy and ambiguous. In crowdsourcing, for example, human annotators may assign conflicting class labels to the same instances. Partial-label learning (PLL) addresses this challenge by training classifiers when each instance is associated with a set of candidate labels, only one of which is correct. While early PLL methods approximate the true label posterior, they are often computationally intensive. Recent deep learning approaches improve scalability but rely on surrogate losses and heuristic label refinement. We introduce a novel probabilistic framework that directly approximates the posterior distribution over true labels using amortized variational inference. Our method employs neural networks to predict variational parameters from input data, enabling efficient inference. This approach combines the expressiveness of deep learning with the rigor of probabilistic modeling, while remaining architecture-agnostic. Theoretical analysis and extensive experiments on synthetic and real-world datasets demonstrate that our method achieves state-of-the-art performance in both accuracy and efficiency.

Cited 0 times · View on doi.org

Continue

Continue Exploring

Ask this paper your own questions, or keep browsing the verified research catalogue.