A question sits at the center of almost every serious debate about artificial intelligence: when a large language model produces an answer, is it actually reasoning, or is it generating text that merely resembles reasoning? The distinction is not academic. It shapes what tasks AI can be trusted to perform, how much human oversight is required, and what the real consequences of deploying these systems will be.
Melanie Mitchell, a cognitive scientist and computer scientist at the Santa Fe Institute, argues that the field currently lacks adequate methods to answer that question with any rigor. Her position is both a critique and a research agenda, and it draws on a surprising source of inspiration: the scientific study of babies and animals.
AI as Alien Intelligence, and Why That Framing Changes Everything
Mitchell describes AI as a form of “alien intelligence.” The phrase is precise, not poetic. These systems have been trained on vast quantities of human-generated language, images, and text from across the internet, yet the way they learn, process information, and produce outputs is fundamentally different from how human cognition works. The surface resemblance to human reasoning is real. The underlying mechanism is not.
This creates a measurement problem. Most current evaluations of AI capability are designed to test performance on tasks, not to probe the cognitive processes behind that performance. A system can score well on a benchmark without doing anything remotely like what a human does when solving the same problem. That gap between output and process is where the hard questions live.
What Mitchell proposes is that AI researchers look to fields that have already grappled with this exact challenge. Developmental psychologists study how infants acquire intelligence without being able to ask them directly. Comparative psychologists study cognition in birds, dogs, and dolphins, none of which can explain their own mental processes. Both fields have developed experimental methodologies for inferring what is happening inside a mind from the outside. Mitchell argues those methods can and should be adapted to study AI systems.
The Clever Hans Problem: A Century-Old Warning
One of the most instructive examples Mitchell raises comes from the early 1900s: a horse named Clever Hans, famous at the time for apparently being able to perform mathematical calculations. Audiences were astonished. Investigators eventually discovered that the horse was not doing arithmetic at all. It was reading subtle, involuntary physical cues from the humans around it, cues that signaled when to stop tapping its hoof. The horse was genuinely responding to something, just not to the thing everyone assumed.
The parallel to AI evaluation is direct. When a language model produces a correct answer, observers tend to attribute that success to the cognitive process they would use to reach the same answer. But the model may be responding to entirely different signals in the input, patterns in its training data, statistical regularities that have nothing to do with the reasoning chain the output appears to describe. Without careful experimental design, it is easy to be fooled in the same way audiences were fooled by Clever Hans.
This is not a hypothetical concern. It is the core methodological challenge Mitchell and others in cognitive science are urging the AI field to take seriously. Mike Frank at Stanford has written about how AI researchers could draw on developmental psychology for exactly this reason, and Mitchell extends that argument to animal cognition as well. The field has been thinking about how to study minds it cannot directly interrogate for at least a hundred years, as Mitchell notes. That accumulated knowledge is largely untapped in mainstream AI research.
Why This Matters Beyond the Lab
The stakes here extend well past academic methodology. If the tools used to evaluate AI systems are systematically misleading, then decisions made on the basis of those evaluations, about deployment, about trust, about the level of human oversight required, rest on shaky ground.
Mitchell points out that the AI community itself is polarized along this axis. Some believe current systems have surpassed human intelligence in meaningful ways. Others believe they remain far from anything resembling human-like cognition. That disagreement is not just a matter of opinion. It reflects a genuine absence of shared, rigorous methods for measuring what these systems actually do.
This is what most coverage of AI capability misses. The debate is not primarily about whether AI is impressive. It clearly is. The debate is about whether the field has the scientific tools to know what it is measuring when it measures AI performance. Mitchell’s argument is that it does not, not yet, and that building those tools requires borrowing from disciplines that have been quietly solving adjacent problems for a very long time.
AI systems are already being used to assist with consequential tasks, including, as Mitchell and her interviewers note, recent breakthroughs in mathematics. The question of what these systems are actually doing when they succeed is not a philosophical luxury. It is a practical prerequisite for knowing when to trust them.
In Short
Evaluating AI intelligence requires more than testing whether a system gets the right answer. Cognitive science and developmental psychology offer experimental frameworks for probing what is actually happening inside a mind, frameworks that the AI field has largely ignored. Without better measurement methods, the field risks repeating the mistake of Clever Hans: mistaking a convincing performance for the real thing.
Based on reporting from Quanta Magazine.