Artificial intelligence systems can write fluently, reason through complex problems, and generate responses that often feel remarkably human. Yet a recent study published in PNAS Nexus suggests that one of the most basic cognitive skills humans exercise daily, staying focused on a goal when distractions compete for attention, may expose a fundamental gap in how today’s AI systems actually work.
The research, led by Suketu Patel, used a decades-old psychological experiment to probe that gap. What it found is worth understanding carefully, because it reframes a question that often gets lost in discussions about AI capability: the difference between producing impressive outputs and genuinely controlling attention.
A Test Designed to Stress Human Focus
The Stroop task is one of the most well-known tools in cognitive psychology. The setup is straightforward. Color words such as “red,” “blue,” or “green” are displayed in colored ink. Sometimes the word and the ink color match. Sometimes they conflict, as when the word “red” appears printed in blue ink. The participant’s job is to name the ink color, not read the word.
This sounds trivial. It is not. Reading words is an automatic habit for most people, something the brain does without conscious effort. Naming the ink color requires actively suppressing that habit, which is precisely what makes the task useful for measuring executive control: the mental processes that allow people to regulate attention, resist competing impulses, and stay oriented toward a specific goal.
Psychologists have used the Stroop task for decades because it isolates something real and measurable about how the brain manages conflict between automatic responses and deliberate intentions. Patel’s research team applied the same logic to large language models.
What Happened When AI Took the Test
The researchers tested several leading AI systems, including GPT-4o, GPT-5, Claude 3.5 Sonnet, Claude Opus 4.1, and Gemini 2.5, on Stroop-style lists of varying lengths.
With short lists of five color words, the models generally performed well. GPT-4o reached 91% accuracy at that length. The picture changed as lists grew longer. At ten words, GPT-4o’s accuracy fell to 57%. At forty words, it dropped to 15%. Claude 3.5 Sonnet held stable performance through lists of twenty words before experiencing a sharp decline, falling to 24% accuracy at forty words.
The situation deteriorated further when matching and mismatched items appeared together in the same list. Under those conditions, accuracy for the mismatched items dropped to nearly zero in some cases. The models were not simply making random errors. They were defaulting to reading the words rather than identifying the ink colors, which is exactly the behavior the task is designed to suppress.
This is what most coverage of AI capability misses. The failure is not about knowledge or language skill. These models know what ink colors are. They understand the instructions. The problem is that as the task becomes more demanding, they lose the ability to consistently override the response they were most heavily trained to produce: processing and reproducing text.
Why This Matters Beyond the Lab
The contrast with human performance is instructive. People also find mismatched items harder than matching ones, and reading words is a stronger habit than naming colors. Yet most individuals can maintain high accuracy and stable performance even across long lists of conflicting stimuli. The human brain has mechanisms for sustaining goal-directed attention over extended sequences, even when automatic responses are pulling in a different direction.
Current AI models, the study suggests, do not appear to have an equivalent mechanism. Their performance does not degrade gradually as cognitive load increases. It collapses.
This distinction matters for anyone thinking seriously about where AI fits into real workflows. Tasks that are short, well-defined, and low in competing signals tend to go well. Tasks that require sustained attention across longer sequences, or that involve resisting the most statistically likely response in favor of a specific instruction, are where the architecture shows its limits.
The findings do not suggest that AI systems are useless or that their capabilities are illusory. They suggest something more precise: that the mechanisms underlying AI performance are genuinely different from the mechanisms underlying human attention, and that those differences become visible under specific conditions. Understanding where those conditions arise is more useful than either dismissing the technology or overstating what it can do.
AI tools work best when they augment human judgment rather than replace it, particularly in contexts where sustained focus, resistance to distraction, and consistent adherence to a specific goal over time are what the task actually demands.
In Short
A research team led by Suketu Patel used the Stroop task, a classic psychology experiment measuring attention and executive control, to test leading AI models including GPT-4o, Claude 3.5 Sonnet, and Gemini 2.5. Performance held up on short lists but collapsed as lists grew longer, with GPT-4o dropping from 91% accuracy at five words to 15% at forty. The models increasingly defaulted to reading words rather than following the instruction to name ink colors. Humans, despite facing the same automatic bias toward reading, maintain stable performance under the same conditions. The study points to a real architectural difference between human attention and how current large language models process information, one that becomes consequential in any task requiring sustained, goal-directed focus over extended sequences.
Based on reporting from ScienceDaily AI.