On September 3 of last year, Greg Brockman, president of OpenAI, announced the release of GPT-6 Astra, a new artificial intelligence model, and declared, "Welcome to the era of AGI." AGI refers to artificial general intelligence, a theoretical form of AI that could match human capabilities across all domains. Just three days later, Jensen Huang, CEO of Nvidia, a company that supplies the powerful chips used to train these models, wrote on X, "AGI has arrived." However, Sam Altman, CEO of OpenAI, still considered the term "not very useful" as of summer 2025, highlighting the ongoing debate around what AGI truly means.
The timing of OpenAI's announcement may have been influenced by a long-standing contract with Microsoft, which had been in place since 2019. This agreement made AGI a critical milestone, as the day OpenAI achieved AGI, Microsoft would lose some of its rights to the company's technologies. The definition of AGI had major financial implications and was revised multiple times during negotiations, from a vague formula in 2019 to a threshold of $100 billion in profits, as reported by the media outlet The Information at the end of 2024, and later to validation by a panel of independent experts in October 2025. On April 27, 2026, the clause related to AGI was removed from the contract, and the term was officially announced 129 days later. This timing became significant as OpenAI and its competitor Anthropic had filed confidential documents for an initial public offering in June, and generating buzz with a low-cost announcement was valuable during this period.
Nvidia, which provides the chips used to train Astra and is also a shareholder in OpenAI, is not a neutral observer in this discussion. The definition of AGI is influenced by the company that supplies the hardware, much like a thermometer that is calibrated by the seller of the heating system. Human resources directors, governments preparing regulations, and employees contemplating their job security all rely on a measurement system created and named by the seller. AGI has thus become the only category in the sector without a recognized arbiter, making it the only one where OpenAI cannot lose.
Despite these claims, the performance of Astra, as measured by various tests, raises questions about its capabilities. ARC Prize, the foundation managing ARC-AGI-3, one of the most demanding reasoning tests in the field, measured Astra in two different ways. It achieved a 99.9% success rate using a special device designed to allow the model to retain its reasoning from one question to the next, but only 62.7% using the standard device common to all models. The same model taking the same exam thus differed by 37 points depending on its surroundings, and the foundation itself noted that "passing its test would not constitute proof of AGI." On FrontierMath, another mathematics test highlighted at the launch, OpenAI funded the design of the test and had access to most of its problems and solutions, which is akin to taking an exam where one can read part of the subject.
Since no one agrees on the definition of AGI or how to measure it, the results depend as much on the testing device as on the intelligence of the model, and this gap has a very concrete translation in business. Two teams with the same subscription to an AI tool rarely achieve the same results, the difference not coming from the tool but from how processes are described, how data is prepared, and how employees know how to formulate a request and verify the response. What the laboratories call the device around the model, a company calls its organization, and an organization is built through training.
Astra remains an excellent model, and this good news needed no exaggeration. OpenAI attributes unprecedented results in mathematics on gaps between prime numbers, which remain to be validated by the scientific community, and its performance places it at the level of Fable 5.1, one of the most advanced models of Anthropic. The race continues without any indication that large language models, the technology that powers these tools, have reached their peak.
However, Astra only achieves 41.4% on AutomationBench, a test that measures the chaining of several steps in professional applications, meaning that a model capable of advancing mathematical research still fails more than half the time to chain the actions that an administrative assistant performs every day. This is the true portrait of current AI, brilliant in some areas, fragile in others, and very far from an intelligence capable of acting in all areas and understanding the world around it.
A majority of researchers, however, doubt that large language models are sufficient to achieve this, as 76% of the 475 researchers surveyed in March 2025 by the American Association for Artificial Intelligence (AAAI) considered it improbable that the simple increase in current approaches would lead to this. Attention is now turning to other architectures, such as "world models" intended to represent the functioning of the real world.
What disappears first, therefore, are not jobs but the gestures that link two software applications, such as copying data from a spreadsheet to a customer file, reissuing an invoice, or compiling three exports into a dashboard, tasks that do not appear on any job description but fill entire days. Automating them frees up time for what the machine cannot do, provided that employees themselves know how to design these automations. Rather than debating an AGI about which no one shares a definition, Europe would benefit from requiring independent criteria to measure what AI can actually do and from mass training its employees and job seekers to use it, because AGI can wait, but the rise in skills cannot.
OpenAI's GPT-6 Astra and the Debate Over AGI Definitions
AI-rewritten from original reportingHow it works
agiopenainvidiaai-benchmarksworkforce-impactagreement



