The Brief
Science & Discovery 5 min read

From 246 to 186: What AI Just Did to a Human Math Record

NAVION

Share

A mathematician spent two years refining a technique that had been stalled for over a decade. She moved a famous mathematical boundary by six units. Three days later, her record was gone.

That sequence of events, compressed into the first week of September, has become a flashpoint in a debate that extends well beyond number theory. It raises questions about how knowledge is created, who gets credit for it, and what happens to the culture of a field when machines can accelerate past human effort almost instantly.

A Twelve-Year Stalemate, Then a Six-Unit Leap

The twin prime conjecture is one of mathematics’ most enduring open problems. It asks whether there are infinitely many pairs of prime numbers separated by exactly two, such as 3 and 5, or 17 and 19. The conjecture is widely believed to be true, but no one has proven it.

In 2013, mathematician Yitang Zhang, now at Sun Yat-sen University in Guangzhou, China, made a landmark contribution. He proved that some gap size smaller than 70 million must repeat infinitely among prime numbers. That was the first concrete upper bound on the problem. Collaborative efforts quickly drove that number down to 246 by mid-2014. Then progress stopped. For twelve years, 246 was the record.

Julia Stadlmann, a postdoctoral researcher at the University of Illinois Urbana-Champaign, began working intensively on the problem roughly two years ago, after completing her doctorate under James Maynard at the University of Oxford. Maynard had helped set the previous record and won the Fields Medal in 2022, partly for his work on prime gaps. Stadlmann’s task was to integrate more recent mathematical techniques with Zhang’s earlier framework, using a method called sieve theory, which filters out composite numbers using carefully calibrated weights. Determining the optimal weights required calculating volumes in high-dimensional space with only an approximate picture of what those regions look like. Andrew Granville, a number theorist at the University of Montreal, described the difficulty plainly: it goes beyond finding a needle in a haystack.

By early summer 2026, Stadlmann had moved the bound from 246 to 240. It was the first advance in more than a decade. Kannan Soundararajan of Stanford University called it “really quite impressive,” citing her persistence and courage.

Three Days at the Top

Stadlmann’s postdoctoral mentor, Kevin Ford, urged her to post her result immediately to the math preprint archive when rumors surfaced that OpenAI was working on the same problem. She did. The record lasted three days.

Researchers at Axiom Math, an AI startup that had been verifying earlier proofs from 2013 and 2014, pivoted quickly after Stadlmann’s result appeared. They incorporated her new ideas into a theorem-proving AI and, working through the night, pushed the bound down to 212.

Within two hours of Axiom’s announcement, OpenAI published its own result. The company had been working on a related problem, the largest gaps between primes, as a benchmark for its newest large language model, GPT-6 Astra. The small-gaps problem was, according to OpenAI computer scientist Sébastien Bubeck, something of an add-on pursued by mathematicians within the company. GPT-6 Astra brought the bound down to 186.

OpenAI says it was in contact with Maynard during this period but did not reach out to Stadlmann directly. That omission drew sharp criticism. In mathematics, the professional norm is to step back when another researcher, especially a junior one, is already working on a problem, or to reach out and collaborate. Granville stated directly that he does not see evidence OpenAI carefully considered how it treated her. Stadlmann herself says she does not feel mistreated. Many of her colleagues disagree on her behalf.

What Gets Lost When Machines Race Ahead

This is what most coverage of the story misses: the dispute is not really about who holds a record. It is about what mathematical progress is for.

Terence Tao of UCLA, a Fields Medalist and one of the most widely read voices in mathematics, wrote that what unfolded reflects deliberate choices to abandon any pretense of gaining human understanding, using AI tools solely to achieve a benchmark. That framing points to a genuine tension. When a machine produces a proof, it may be formally correct and yet opaque. Mathematicians refer to outputs filled with content that humans cannot immediately parse as “AI slop.” A valid proof that no one can understand does not advance the field in the same way a human-constructed proof does.

The deeper concern is about what the struggle itself produces. Tao, who collaborated on earlier small prime-gap efforts, has noted that for problems like these, the value lies not in the answer but in what emerges from the process: new techniques, new ways of framing problems, a richer understanding of mathematics as a whole. Granville offered a concrete illustration of AI’s genuine usefulness, noting that a proof that would have taken him a month now takes two hours. The tool is powerful. The question is whether deploying it to race past human researchers, without coordination or transparency, serves the broader goal of building knowledge, or simply accumulates benchmarks.

Bubeck stated that OpenAI does not intend to push further on the prime-gaps problem and that the company’s strategy is to empower mathematicians. Whether that framing holds as AI capabilities grow is a question the field is now actively wrestling with.

In Short

A human mathematician broke a twelve-year-old record. Within days, two AI systems broke hers. The numbers, 246, 240, 212, 186, tell one story. The more important story is about professional norms, the opacity of machine-generated proofs, and whether the race to capture mathematical benchmarks is compatible with the slower, richer process of actually understanding mathematics.

Based on reporting from Science News.

Written by

NAVION