Engineering · MapleScholar Plus

Breaking the Memory Wall: ThAME Heterogeneous Accelerator for Mixture-of-Experts LLMs

Mixture-of-Experts (MoE) AI architectures require massive memory bandwidth to activate sparse expert networks, bottlenecking conventional GPU clusters; ThAME integrates monolithic 3D memory to deliver high-throughput, energy-efficient MoE inference.

Author
Pratyush Dhingra et al.
Published
2026
Journal
arXiv (Cornell University)
Last updated
September 2026
Breaking the Memory Wall: ThAME Heterogeneous Accelerator for Mixture-of-Experts LLMs

Frontier language models are increasingly adopting Mixture-of-Experts (MoE) architectures, where different neural network 'experts' are dynamically activated for specific tokens, enabling trillions of parameters with sparse computation.

However, serving MoE models in production triggers the infamous 'memory wall': GPUs spend more energy moving billions of expert weights across off-chip memory buses than actually performing matrix multiplications.

ThAME (3D Memory-Enabled Heterogeneous Accelerator) solves this memory bottleneck by stacking high-density 3D memory layers directly on top of specialized compute logic. The close physical vertical integration delivers massive memory bandwidth and near-zero interconnect latency tailored for dynamic MoE expert routing.

ThAME slashes generative AI inference energy consumption while accelerating token generation speeds by multiples, offering an architectural blueprint for serving trillions-parameter AI models sustainably.

Reference

Dhingra, P., Pal, P. K., Doppa, J. R., & Pande, P. P. (2026). ThAME: 3D Memory-Enabled Heterogeneous Accelerator for LLM Mixture of Experts (Version 2). arXiv.

Title

ThAME: 3D Memory-Enabled Heterogeneous Accelerator for LLM Mixture of Experts

Abstract

Mixture of Experts (MoE) architectures have emerged as a dominant paradigm for scaling Large Language Models (LLMs). However, MoE inference on conventional hardware is constrained by three fundamental bottlenecks. These encompass the massive memory bandwidth required to fetch non-contiguous expert weights, the non-deterministic scatter-gather traffic generated by input-dependent token routing, and the tail-latency dependency imposed by synchronous expert output aggregation. To address these challenges, we propose ThAME, a three-dimensional (3D) heterogeneous multi-chiplet architecture for MoE inference. ThAME employs Ferroelectric Field-Effect Transistor (FeFET)-based non-volatile and DRAM-based volatile memory chiplets with a co-designed compute mapping strategy that aligns the distinct computational profiles of attention mechanisms and expert routing. Furthermore, we design a specialized Network-on-Chip communication backbone optimized to mitigate the bottlenecks associated with non-deterministic token routing traffic across the combinatorial space of input-dependent MoE traffic patterns. Experimental results demonstrate that ThAME outperforms state-of-the-art counterparts by up to 15.7x in terms of speedup and improves energy efficiency by up to 9.8x.

Cited 0 times · View on doi.org

Continue

Continue Exploring

Ask this paper your own questions, or keep browsing the verified research catalogue.