Chat
Neuroscience · MapleScholar Plus

The Distraction Trap: Why the World's Best AI Flunks a 100-Year-Old Psychology Test

Leading neural networks achieve human-level scores on short academic exams; they suffer an eighty-percent accuracy collapse when forced to resolve repetitive mental conflicts across long tasks. By testing state-of-the-art language models with classic psychology tasks, cognitive scientists proved that transformer attention lacks the executive mental bouncer required for sustained concentration.

Author
S. M. Patel et al.
Published
2026
Journal
PNAS Nexus
Last updated
September 2026
The Distraction Trap: Why the World's Best AI Flunks a 100-Year-Old Psychology Test

In tech boardrooms and marketing demos, modern generative AI is hailed as approaching human-like reasoning. However, when deployed on long, repetitive legal audits or complex software refactoring, these systems mysteriously lose track of basic rules and produce glaring logical errors.

Cognitive scientists exposed this weakness using the classic psychology color-word test: naming the ink color of the word "RED" written in blue ink. While human brains easily filter out the distraction using executive mental focus, transformer neural networks become overwhelmed as the list grows, collapsing from ninety percent accuracy down to fifteen percent.

This discovery proves that making models bigger will not solve reasoning failures. By pinpointing the missing executive control circuits, by guarding against premature automation in high-stakes law and medicine, and by inspiring biologically grounded neural architectures, cognitive testing reshapes AI engineering.

Reference

Patel, S. C., Wang, H., & Fan, J. (2026). Deficient executive control in transformer attention. PNAS Nexus, 5(6).

Title

Deficient executive control in transformer attention

Abstract

Although transformers in the large language models (LLMs) effectively implement a self- attention mechanism that has revolutionized natural language processing, they lack an explicit implementation of executive control of attention found in humans, which is essential for resolving conflicts and selecting relevant information in the presence of competing stimuli, and is critical for adaptive behavior. To investigate this limitation in LLMs, we employed the classic color Stroop task that is widely regarded as the gold standard for testing executive control of attention. Our results revealed a typical conflict effect of better performance in terms of accuracy in the congruent condition (e.g., naming the ink color of the word RED in red) compared to the incongruent condition (e.g., naming the ink color of the word RED in blue), which is similar to human performance in short sequences. However, as sequence length increased, the performance degraded toward chance levels on the incongruent trials despite maintaining excellent performance on congruent trials and near-perfect word reading ability. These findings demonstrate that while transformer attention mechanisms can achieve human-comparable performance in smaller contexts, they are fundamentally limited in their capacity for conflict resolution across extended contexts. This study suggests that incorporating executive control mechanisms akin to those in biological attention could be crucial for achieving more general reasoning and reliable performance toward artificial general intelligence.

Cited 2 times · View on doi.org

Continue

Continue Exploring

Ask this paper your own questions, or keep browsing the verified research catalogue.