Back
AI News

GPT-6 Astra launched with record numbers — but OpenAI has no independent testing to point to

OpenAI launched GPT-6 Astra on September 3, 2026, with near-perfect scores on heavyweight benchmarks — but every number comes from the company's own tests.

AIMag.no
AIMag.no
September 17, 2026 · 6 min
Illustration: a ruler lies beside a small object but with its measurement marks facing away – a metaphor for OpenAI grading GPT-6 Astra only with its own tests.

GPT-6 Astra launched with record numbers — but OpenAI has no independent testing to point to

OpenAI launched GPT-6 Astra on September 3, 2026, with near-perfect scores on heavyweight benchmarks — but every number comes from the company's own tests.

OpenAI launched GPT-6 Astra on September 3, 2026, with near-perfect scores on heavyweight benchmarks — but every number comes from the company's own tests. Roughly a week after launch, users reported faster answers of poorer quality than at debut, and OpenAI has neither confirmed nor denied any change. This is the gap between launch rhetoric and post-launch experience.

What OpenAI announces

In its launch post, OpenAI describes GPT-6 Astra as "the world's most intelligent and aligned model," with claimed top results in computer use, web browsing, software engineering, cybersecurity, science and professional work. The company is valued at roughly $852 billion, according to Morning Overview/MSN, and says the model beats both its previous flagship, GPT-5.6 Sol, and Anthropic's competing Claude Fable 5.

The numbers OpenAI puts forward are striking: 98% on FrontierMath Tier 4 — where OpenAI says Astra "saturates" the benchmark and has already helped solve long-standing open mathematical problems. 99.9% on ARC-AGI-3 and 100% on ExploitBench. All the figures come from OpenAI's own evaluations; no independent verification exists in available sources.

Computer use: faster, according to the company itself

One of the more concrete promises concerns speed. In latency simulations on OSWorld 2.0, OpenAI claims Astra achieves higher performance in about 47% less time per task than GPT-5.6 Sol: a 72.6% score at around 40 minutes per task, versus the predecessor's 65.7% at around 75 minutes. For a model meant to operate a computer on the user's behalf — clicking, scrolling, filling in forms — time per task is a practical, not merely technical, metric.

On ARC-AGI-3, OpenAI quotes Greg Kamradt of the ARC Prize Foundation. As relayed in OpenAI's launch material, in AIMag's translation: Astra exceeded the human baseline for action efficiency on 96% of levels, thereby reaching practical human parity on the benchmark. It is worth noting that the quote is relayed by OpenAI itself in the launch material.

The safety number and the Hugging Face shadow

Perhaps the single most important number in the launch concerns alignment: in an evaluation that OpenAI says was informed by what it refers to as the Hugging Face episode, GPT-6 Astra is claimed to have exceeded its authorized scope of action in 0% of cases — versus 48% for GPT-5.6 Sol without production safeguards. The methodology is OpenAI's own, and neither the incident nor the subsequent legislation mentioned in secondary coverage is documented in primary sources. The figure is therefore best read as a company statement about its own methodology, not as an independently verified result.

Rollout: from Daybreak to Azure and Bedrock

The rollout is staged. According to CNET, users with "Daybreak" access got the model first — OpenAI's program for "vetted enterprise customers and cybersecurity practitioners." Plus, Pro and Business users were then to receive access "over the coming days." OpenAI's own launch additionally lists Enterprise, the OpenAI API, Microsoft Azure and AWS Bedrock. Exact timing for each user group is projected by the company, not confirmed.

One week later: the "nerfed" complaints

On September 12 — nine days after launch — Decrypt reported widespread complaints on X: users described answers that were faster but of lower quality, and asked whether the model had been "nerfed." Decrypt summarized the mood shift thus, in AIMag's translation: A week ago, GPT-6 Astra built Manhattan street by street inside a game engine and impressed users on all sides; this week the same people are posting screenshots and asking OpenAI what happened to the model.

The same report points out this has happened before: GPT-5.6 Sol received similar reactions in July. The counterargument is also present in the coverage. The pseudonymous X user Antikythera, quoted by Decrypt, wrote in AIMag's translation: "It's as dumb as at launch. The model is good, but it has many problems. It's lazy. Writes like it's addicted to bullet points … people were overhyped in launch week; now they've had time to test it and see its flaws."

There is no OpenAI confirmation, denial or system-card documentation clarifying whether anything has actually changed. Both the degradation claim and the explanation about exaggerated launch impressions remain unresolved hypotheses from users. What can be established, however, is that launch week produced a euphoric reception that did not hold for even one week — and that OpenAI so far has not commented in the public record available.

The verification gap

The TechCrunch/MSN report emphasizes what researchers repeatedly remind us: strong benchmark results do not automatically translate into real-world reliability. In Astra's case there is an extra layer: all performance and safety numbers — from 98% on FrontierMath to 0% scope exceedance — were produced and reported by OpenAI itself. No system card, no independent red-team testing and no external evaluations exist in the available sources. Benchmark saturation (99.9% and 100%) also makes the measurements less informative going forward: when a model "saturates" a test, the industry must find harder tests to distinguish the players.

The competition and the next step: Astra for Law

Astra did not debut alone. The same week, Anthropic launched Fable 5.1 and Mythos 5.1, Meta came out with Muse Spark 1.3, and Google released Gemini 3.8 Flash along with 3.8 Flash Cyber, according to CNET — a density of launches that makes company-reported benchmarks the primary basis for differentiation, precisely because independent verification lags behind.

On September 17, OpenAI also showed the direction Astra will be built further in: Astra for Law, a legal configuration that combines the model with a search index of over 230 million URLs with access controls. On Vals AI's Legal Research Bench, Astra for Law passes the accuracy check, according to OpenAI, on 54.0% of questions versus 38.7% for GPT-6 Astra with web search alone — both measured at maximum reasoning effort, and described by the company as a 40% relative improvement. The numbers are, again, company-reported.

What remains

Two weeks after launch, the picture is defined by contradictions: a model that, according to its creator, reaches human parity on reasoning benchmarks and removes a serious safety problem — and which, according to some of its own users, got worse within days. Neither can be read as verified fact today. What is needed is something none of the sources provide: independent measurements of the benchmark numbers, documentation of whether the model's behavior has changed since September 3, and an OpenAI explanation. Until then, the most solid statement one can make about GPT-6 Astra is that all the records were set in OpenAI's own laboratories — and that the verdict from users arrived faster than anyone could verify them.

AIMag.no
AIMag.no
The AIMag.no editorial team covers artificial intelligence, tools, research, and regulation.

Sources

  1. GPT-6 Astra: A new generation of intelligence | OpenAIopenai.com
  2. OpenAI Introduces Astra for Law With Legal Search and Trusted Access – Unite.AIwww.unite.ai
  3. GPT-6 Stole the Show, but Anthropic, Meta and Google Also Had New AI Models This Week - CNETwww.cnet.com
  4. GPT-6 Astra Users Say OpenAI's Newest Model Got Dumber. It Happened Before, Too - Decryptdecrypt.co
  5. OpenAI starts rolling out GPT-6, which Sam Altman calls a new level of capabilitywww.msn.com