The Brief
How AI Works 5 min read

100,000 Years of Language, and Kids Still Win

NAVION

Share

For at least 100,000 years, only one kind of entity on Earth could learn a human language to full fluency: a human child. That changed roughly four years ago. Large language models can now hold conversations that are, on the surface, indistinguishable from human exchange. And yet, underneath that fluency, something deeply strange is happening. These systems require a staggering quantity of data to reach a level of competence that a toddler achieves almost effortlessly. Understanding why that gap exists is now one of the more consequential open questions in both cognitive science and AI research.

The Scale of the Problem Is Hard to Grasp

The numbers involved are difficult to hold in the mind. A preteen raised in a linguistically rich environment may have heard roughly 100 million words by the time they reach adolescence. Add reading to the picture, and that figure might climb to around 300 million words by age 20. Meta’s open-weight model Llama 3.1, by contrast, was trained on 15 trillion tokens, where tokens are word-like units of language. Frontier models may be training on ten times that amount, according to Ethan Gotlieb Wilcox, a cognitive scientist and linguist at Georgetown University.

To make the contrast tangible: if you printed all the words used to train a modern large language model, the stack of paper would reach past the International Space Station. The 100 million words a preteen has heard would stack up to about 20 meters. Wilcox puts it another way: Claude has seen the amount of language that an entire city will experience in one generation.

This is what researchers call the data efficiency gap. Children learn language from a tiny fraction of the input that machines require, and they do it faster, more reliably, and with far less explicit instruction. Toddlers typically begin producing grammatically correct sentences after hearing somewhere between 10 and 30 million words. Training GPT-2 on 30 million words, as Michael C. Frank, a cognitive scientist at Stanford University, points out, produces a nonsense generator. It does not produce anything resembling a child.

A Debate That Has Shaped Both Linguistics and AI

The question of how children acquire language is not new, and the competing answers to it have had real consequences for how AI developed. In the 1950s, the MIT linguist Noam Chomsky argued that children are born with hardwired grammatical knowledge. His reasoning was that language, especially its syntax, is too complex and children’s exposure to it too limited for learning to happen purely through experience. He called this the “poverty of the stimulus.” Against him stood the psychologist B.F. Skinner, who held that language was learned entirely through environmental conditioning, the same way animals learn to respond to rewards.

Chomsky’s view dominated American linguistics for decades, and it shaped early AI research directly. Researchers tried to teach computers language by coding grammatical rules explicitly into programs, a rule-based approach that became part of what was called symbolic AI. That approach largely failed to produce systems capable of handling real human language at scale. Interest in the field cooled significantly during the period known as the AI winter, which began in the 1970s.

The comeback came through a different route entirely. Neural networks, which learn by recognizing statistical patterns rather than following explicit rules, began outperforming rule-based systems as hardware became cheaper and the internet provided vast quantities of text. By 2018 and 2019, models like BERT and GPT-2, built on a new architecture called the transformer, demonstrated that learning from massive amounts of data could work for language. ChatGPT’s success in 2022 made that visible to the general public.

The irony is pointed. Large language models are, in essence, exactly the kind of statistical learning machines that Chomsky argued could never acquire language. They have no innate grammar, no hardwired rules, no biological architecture. And yet they work, at least at scale. Children, meanwhile, manage the same feat with a fraction of the data and without any of the computational infrastructure. Neither side of the old debate fully explains either case.

Why This Question Matters Beyond the Lab

There is a practical urgency here that goes beyond academic debate. The internet is not infinite. Researchers suggest that the supply of easily available text data for training could run dry as early as the 2030s. If AI systems can only improve by consuming more data, that trajectory has a ceiling. Children demonstrate that the ceiling is not fundamental. It is a limitation of current methods, not of the task itself.

Closing the data efficiency gap could open possibilities that current models cannot reach: AI systems trained effectively on video rather than text, language tools built for minority language communities where large text corpora simply do not exist, and models that generalize from limited examples the way humans do. Studying how children learn also offers a way to test longstanding hypotheses about the human mind, including whether language acquisition requires innate biological structures or whether the right kind of learning environment is sufficient.

What this field is really asking is a version of a much older question: what does it mean to understand language, and how much of that understanding is uniquely human? AI has not answered that question. It has made it more urgent.

In Short

Children learn language from a fraction of the data that AI systems require, a gap so large it can only be described through analogy. Researchers are studying how children pull this off because the answer could reshape how AI models are built, and because the supply of training data has limits. The debate between innate knowledge and statistical learning, which shaped both linguistics and early AI, is still unresolved. Large language models are powerful evidence that statistics can go a long way. Children are equally powerful evidence that statistics alone are not the whole story.

Based on reporting from MIT Technology Review.

Written by

NAVION