The Brief
How AI Works 5 min read

AGI Is "Here." The Definition Just Keeps Moving

NAVION

Share

OpenAI president Greg Brockman recently declared that the “AGI era” has begun, following the launch of GPT-6 Astra. Nvidia CEO Jensen Huang echoed the announcement publicly. The claim landed with considerable weight: artificial general intelligence, long treated as the distant horizon of AI research, was being presented as a present-day reality. The problem is that almost no one outside those companies agrees on what that reality actually means.

A Benchmark Score That Changes With the Software

Astra’s most cited evidence for AGI-level capability is its performance on ARC-AGI-3, a benchmark designed by AI researcher François Chollet specifically to resist the kind of targeted training that has made many AI tests unreliable. The benchmark places models in unfamiliar video game environments and measures whether they can figure out the rules, adapt, and win. It is, by design, harder to game than most.

Astra performed remarkably well. Using OpenAI’s own testing harness, the software layer that mediates between the model and the benchmark, it scored 99.9%. Using ARC-AGI-3’s standard testing setup, designed to give all models a more uniform interface, the score dropped to 62.7%. That gap is not a minor technical footnote. It raises a direct question: is the model genuinely capable, or is it partly optimized for the conditions under which it is being evaluated?

Chollet himself addressed this directly. The benchmark, he explained, tests a non-exhaustive set of attributes at small scales. Real-world intelligence involves far longer time horizons for learning, far greater complexity in modeling the world, and far more ambiguous goals. Solving ARC-AGI-3, in his view, is a meaningful sign of progress. It is not proof of AGI, and it was never intended to be.

The Definition Problem Is the Real Story

Here is what most coverage of this announcement misses: the debate about whether AGI has arrived is inseparable from the debate about what AGI even means. There is no universally accepted definition, and that absence creates significant room for interpretation.

OpenAI’s current definition describes AGI as “highly autonomous systems that outperform humans at most economically valuable work.” The Center for AI Safety uses a more demanding framework requiring high performance across ten distinct cognitive domains, spanning reasoning, memory, and perception. Independent researchers often work from definitions that are stricter still.

Gary Marcus, a professor at New York University and a prominent AI skeptic, described the AGI claim as marketing, arguing that it either misunderstands the original definitions of the term or deliberately lowers the bar. His point about Astra’s performance on the Epoch Capabilities Index is instructive: Astra did set a new record, scoring 169 against a previous high of 163, but Epoch AI’s own analysis found that the improvement remained consistent with the existing trajectory of AI progress. A new record, yes. A discontinuous leap, no.

Ben Goertzel, the researcher who coined the term “artificial general intelligence” in 2005, offered perhaps the most precise framing. After using Astra, his assessment was that the model is superhuman in certain areas, particularly mathematics and programming, while remaining weaker in others. Comparing it to a human being, he said, “ends up being complicated rather than a simple yes or no.” He also identified specific gaps that software scaffolding cannot close: how the model understands itself, how it retains and applies memory across its experience, and how it coordinates competing goals.

Andy Konwinski, who cofounded Databricks, Perplexity, and Laude, put it more plainly. These systems cannot yet think independently over long periods. They can write complex software and identify bugs, but most of the world’s value does not come from software engineering. They are not growing food, building infrastructure, or governing societies.

Why the Label Matters Beyond the Lab

The AGI debate might seem like a technical dispute among researchers, but its implications extend well beyond benchmark scores. When influential figures in the industry declare that a threshold has been crossed, it shapes public expectations, policy conversations, investment decisions, and the way people understand the technology they are increasingly living alongside.

Brockman is not the first to make this kind of declaration. Huang made a similar claim earlier in the year. OpenAI CEO Sam Altman has described the current moment as the beginning of the Singularity, a phase in which AI systems have taken over and accelerated the development of new models, with humans largely removed from that loop. Each of these statements carries weight precisely because of who is saying them, and each one arrives without a shared standard against which to measure the claim.

The cellular network analogy is apt: carriers have historically labeled their networks with the next generation marker before the underlying technology fully met the technical standard. The label drives adoption and perception. The gap between the label and the reality closes later, quietly.

AI capabilities are advancing. That much is not in dispute. What remains genuinely unresolved is whether the concept of AGI, as it has been defined and redefined by the people most motivated to reach it, still describes anything precise enough to be meaningful.

In Short

GPT-6 Astra is a capable model, and its benchmark results represent real progress. But the claim that AGI has arrived rests on a definition that the companies making the claim have shaped to fit their own systems. Researchers who coined the term, designed the benchmarks, and study AI independently are not convinced. The most important thing to understand is not whether AGI is here. It is that the answer depends entirely on who gets to define the question.

Based on reporting from Fast Company - Tech.

Written by

NAVION