In February, at the India AI Summit in New Delhi, Demis Hassabis, co-founder of Google DeepMind, proposed a thought experiment that cuts to the heart of what artificial intelligence can and cannot do. Train a large language model on everything known up to 1911, he suggested, then ask it to derive Einstein’s general theory of relativity, published in 1915. If it could, that would be, in his words, “a good test for AGI.” Several research teams have since attempted versions of this experiment. The results are instructive, and not in the way AI optimists might hope.
Why General Relativity Is the Right Benchmark
Einstein’s 1915 theory is not just a famous piece of physics. It is a genuine paradigm shift: a complete reconceptualization of gravity as a deformation of space-time by mass, rather than a force acting at a distance. The theory now underpins cosmology, informs research on black holes and gravitational waves, and guides GPS satellites in practical use every day. It is, in other words, both intellectually radical and empirically consequential.
That combination is precisely what makes it a useful benchmark. Tom Zahavy, a researcher at Google DeepMind, elaborated on the test in a position paper titled “LLMs can’t jump,” posted in January. His argument draws on a distinction philosophers of science have long recognized: the difference between inductive reasoning, which derives general rules from accumulated data, and abductive reasoning, which invents a cause for a singular phenomenon. Einstein’s leap was abductive. Current AI models are, by design, inductive engines. They identify statistical patterns across vast training sets. They do not invent causes. They do not leap.
What Happened When Researchers Actually Tried
Independent AI researcher Michael Hla, based in San Francisco, built a model he calls Machina Mirabilis, trained on data predating 1900, and tested whether it could reconstruct quantum mechanics, special relativity, and general relativity. He provided deliberate prompts: observations about the photoelectric effect, the essence of Einstein’s “elevator” thought experiment. In some cases, the model produced suggestive outputs. Presented with the photoelectric effect, it described light breaking into “distinct impulses,” hinting at the concept of light quanta. But Hla concluded that the model failed in most cases, lacked genuine understanding of the physics it produced, and was often “parroting words that seem plausible” without any strong internal representation of the world to reason from.
Nick Levine, an independent computer scientist also based in San Francisco, encountered a different but equally revealing problem. His team built a vintage model trained only on material available up to 1930, a date chosen because works published that year entered the public domain in the United States at the start of 2026. The goal was to test whether such a model could independently arrive at concepts developed in the 1930s, including Turing machines, Gödel’s incompleteness theorem, and the postulation of neutrinos. The obstacle was not the model’s reasoning. It was the training data itself. Historical texts are riddled with digitization artifacts and archaic language. Worse, the data proved “maddeningly leaky”: the supposedly pre-1930s model would spontaneously and correctly answer questions about events from the 1950s, absorbing knowledge it was never meant to have.
Sendhil Mullainathan, a computer scientist at MIT, approached the problem differently. His team gave an “orbital mechanics” foundation model synthetic data on planetary systems obeying Newtonian mechanics, then asked it to infer the underlying law of gravitation. The model never recovered the true law. Instead, it inferred a different, incorrect law for each planetary system. It was fitting curves, not discovering principles.
What This Reveals About the Nature of Scientific Discovery
Here is what most coverage of AI in science misses: the hardest breakthroughs in the history of knowledge did not come from having more data. They came from having the right idea about sparse or anomalous data. Kepler’s careful observations of planetary motion were limited in number, but they gave Newton the raw material to construct his laws of gravitation and motion. Einstein’s general relativity emerged from a thought experiment about an elevator, not from a database.
Ido Kaminer, a quantum optics specialist at Technion in Haifa and co-author of a preprint titled “Can AI follow in Einstein’s footsteps?”, argues that a relativity-scale breakthrough is not inherently beyond AI’s reach. But it will require rethinking the foundational principles of how current models are built. The architecture that makes today’s LLMs powerful at prediction and pattern-matching is the same architecture that makes them structurally unsuited to the kind of creative, world-model-building reasoning that produces paradigm shifts.
There is nuance here. An OpenAI chatbot this May disproved an 80-year-old conjecture by Hungarian mathematician Paul Erdős, and Mullainathan describes it as a “genuine conceptual advance,” one that required small new abstractions rather than brute-force search. But even that achievement drew on ideas already present in the mathematical literature. It was synthesis, not revelation.
In Short
AI is becoming a powerful tool for science: it finds patterns, accelerates computation, and can now contribute meaningfully to mathematics. What it cannot yet do is reason from almost nothing to something genuinely new. The Einstein test is not just a clever benchmark. It is a precise description of the gap between what current AI does well and what scientific discovery, at its most transformative, actually requires.
Based on reporting from Nature: Machine Learning.