A proof that took one mathematician seven years to construct, and another researcher a projected five years to formalize, has now been completed by a team of AI agents in eleven days. Anthropic’s Claude model has produced a formal version of Fermat’s Last Theorem, one of the most celebrated problems in the history of mathematics. This is not a new proof. It is something arguably more useful: a machine-readable, machine-verified confirmation that the human proof is correct, built from the ground up in a programming language called Lean.
A 350-Year Problem, and What “Formalizing” Actually Means
Fermat’s Last Theorem states that no whole numbers a, b, and c can satisfy the equation aⁿ + bⁿ = cⁿ when n is a whole number greater than 2. Pierre de Fermat posed the puzzle in the 17th century, famously noting in the margin of a textbook that he had found a proof but lacked the space to write it down. That note haunted mathematics for roughly three and a half centuries.
Andrew Wiles finally proved the theorem in 1995, after working on the problem in secret for seven years. His proof was not without difficulty: a flaw was discovered, and it took Wiles and his collaborator Richard Taylor approximately a year to fix it. The episode illustrates a structural vulnerability in traditional mathematical proofs. They are written in natural language and logical notation, checked by human experts, and vulnerable to errors that can hide inside long chains of reasoning.
Formalization addresses this vulnerability directly. It means translating a proof into computer code, specifically into a formal language that a machine can verify step by step. There is no ambiguity in code. Either the logic holds or it does not. A central repository called Mathlib already stores around 2 million lines of formalized mathematics. Anthropic’s new proof adds 13 million lines to that landscape, covering roughly 29,500 intermediate theorems that were necessary to reach the final result. That makes it more than five times the size of all previous Mathlib content combined, and the largest Lean proof ever written.
How the Agents Actually Worked (and Where They Struggled)
Anthropic’s system did not use a single AI model working in isolation. Multiple separate agents were deployed, each assigned to different portions of the problem, tackling smaller chunks of the theorem in parallel. The process ran continuously and autonomously for eleven days.
Human experts were not absent. They provided what Anthropic describes as “high-level instructions” to keep the work on track. This is worth noting: the agents periodically lost track of the project’s overall state and stopped collaborating effectively. The system required external guidance to recover from those breakdowns. What ultimately helped was a tool originally designed for human mathematical collaboration, called Prove2Me. Once the agents began using it to track their work and coordinate on next steps, the project moved forward more reliably.
Kevin Buzzard, a mathematician at Imperial College London who had been leading a five-year project to formalize the same Wiles and Taylor proof, commented on the result. He noted that the proof rests on “no assumptions other than the axioms of mathematics,” and described the work as multi-layered, touching on algebra, harmonic analysis, geometry, and number theory. His assessment is significant: Buzzard was not a bystander. He had organized a conference in London earlier this year bringing together AI experts, computer scientists, and mathematicians to work on exactly this problem. Anthropic’s announcement overtook that effort entirely.
Why This Matters Beyond the Theorem Itself
Here is what most coverage of this story misses. The theorem itself is not the point. Fermat’s Last Theorem was already proven. What Anthropic has demonstrated is something about the nature of mathematical knowledge and how it can be verified, extended, and built upon.
Buzzard put it directly: if AI can now autoformalize a proof of this complexity, then the automatic formalization of the broader modern mathematical literature becomes a realistic near-term goal. That is a significant shift. Mathematics is the foundation of physics, cryptography, computer science, and engineering. A formalized, machine-verifiable body of mathematical knowledge would allow researchers to build on prior results with far greater confidence, catching errors that human review might miss and accelerating the pace at which new results can be established.
The episode also illustrates something important about how AI systems currently operate in high-stakes intellectual work. They are not autonomous in the way the word is sometimes used loosely. They required human oversight, broke down during coordination, and needed tools designed for human collaboration to function effectively. The eleven-day timeline is remarkable. The need for human guidance throughout is equally informative.
AI here is not replacing mathematicians. It is handling a volume and granularity of logical verification that no human team could sustain at this scale, freeing researchers to focus on the higher-level questions of what to prove and why it matters.
In Short
Anthropic’s AI agents have produced the largest formal proof ever written, verifying Fermat’s Last Theorem in machine-readable Lean code across 13 million lines and 29,500 intermediate steps. The result matters less as a mathematical discovery and more as a demonstration that AI can now formalize complex, layered proofs at a scale and speed previously out of reach. The process required human guidance and broke down more than once. That context belongs in the story too.
Based on reporting from New Scientist.