Back
AI News

First Week With GPT-6-Astra: A Big Jump in 3D and Games — but No AGI, According to the Evaluators

GPT-6, codename Astra, has been out for just over a week, and the first two independent evaluations agree on the broad picture: a substantial capability jump over its predecessor, especially on 3D work, games, computer use, and…

AIMag.no
AIMag.no
September 18, 2026 · 6 min
Illustration: an intricate miniature game level made of folded cardboard and cut paper, with one section left unfinished, cut mid-fold.

First Week With GPT-6-Astra: A Big Jump in 3D and Games — but No AGI, According to the Evaluators

GPT-6, codename Astra, has been out for just over a week, and the first two independent evaluations agree on the broad picture: a substantial capability jump over its predecessor, especially on 3D work, games, computer use, and coordinating subagents — but neither evaluator is willing to call the model AGI. Meanwhile, OpenAI's own benchmark numbers remain independently unverified, and the tension between what the model actually delivers and the "AGI era" rhetoric from OpenAI and Nvidia may be the most revealing part of the launch.

One Week, Two Independent Evaluations

Zvi Mowshowitz published his assessment on September 12, 2026, and PCMag followed with a hands-on test by writer Ruben Circelli on September 14. Both conclude that GPT-6-Astra is a clear step forward, but both reject the AGI framing the companies themselves have set up. It is rare for two independent evaluations in launch week to land so precisely on the same overall conclusion, both on the strengths and on the caveats.

Mowshowitz describes Astra as "the best model for what you would broadly call ambitious projects," and suggests it has the highest raw intelligence factor of any model. He writes that "the jump from Sol to Astra is larger than the jump from Fable 5 to Fable 5.1." The name correspondence between "Sol" and GPT-5.6 is, incidentally, an inference based on the context of the text — neither source explicitly confirms it. He highlights that the model is formidable at anything happening in 3D or involving games, and that it also excels at computer use and coordinating subagents — that is, when one model directs multiple sub-processes or agent instances toward a shared goal.

On regular coding, the picture is more nuanced. According to Mowshowitz, the area receives less focus in this generation, and Astra "is no quantum leap there, but of course very good and making progress over Sol." He also notes that Anthropic Fable 5.1 may still be preferable for some use cases, such as mutual back-and-forth dialogue. Astra's strengths thus lie not primarily where its predecessors' strengths were, but in new categories of work.

Circelli confirms the practical picture on the coding side: GPT-6 completes tasks faster, and in testing it caught errors that GPT-5.6 overlooked. But he also documented the opposite — the model introduced a game-ruining bug that took hours to isolate, and it still requires oversight along the way. Despite all the praise, he concludes that GPT-6 is in no way some kind of superintelligent AI meant to usher in a thousand-year machine reign. One practical limitation: Circelli burned through the weekly usage limit in one hour. The details of the limit cannot be fully verified, as PCMag's text is cut off mid-paragraph, but the price and usage restrictions appear to be a real constraint on professional use — PCMag describes the model as significantly more expensive than its predecessor.

The Numbers — and Who Is Behind Them

The most striking number in the launch is OpenAI's reported ARC-AGI-3 result: 99.9 percent for GPT-6, versus 7.8 percent for GPT-5.6 and 30.2 percent for Anthropic Opus 5. These figures come from OpenAI itself, relayed through PCMag, and are not independently verified. The jump is so large that it should in itself prompt caution: a leap from 7.8 to 99.9 percent on a single benchmark is unusual, and PCMag itself notes skepticism about how benchmark results map to the real world — a skepticism Circelli's own testing partially confirms. The model caught some errors, but introduced one serious one.

The same applies to the claim, attributed to OpenAI's engineering lead via PCMag, that GPT-6 at its lowest intelligence setting outperforms GPT-5.6 at its highest. That is a company statement, not a verified finding, and should be read as such. None of the available sources contain primary documentation from OpenAI — no system card, model card, or press release — so all capability and benchmark claims trace back to the company's own statements relayed by secondary sources. That does not mean the numbers are necessarily wrong, but that for now they are marketing, not documentation.

The AGI Debate, Taken Seriously This Time

Where capability is documented, the AGI framing remains contested. Nvidia CEO Jensen Huang and OpenAI president Greg Brockman have both, according to PCMag, characterized the launch as the entrance to the "AGI era." The PCMag reviewer is unsparing: "calling it AGI is pure marketing spin." Mowshowitz positions himself in between in an interesting way: "This is the first time a debate over whether a model 'was AGI' felt non-silly. I don't think it's AGI, and I'd warn against the danger of using that label too early, but I wouldn't laugh at you for disagreeing."

That is a meaningful difference from earlier launches: even those who reject the label find that it can no longer be dismissed out of hand. At the same time, it remains unclear exactly what Huang and Brockman said — PCMag relays the statements secondhand, and the original wording is not available in the sources. Both Huang and Brockman have a commercial interest in such an expansive characterization, which strengthens the case for treating the claims with distance.

What Remains Unresolved

Perhaps the most intriguing detail is OpenAI's own signaling about more capacity to come. roon at OpenAI wrote on September 3, 2026 that "I have not come close to discovering the limits of what Astra can do." According to Mowshowitz's account, roon also signaled an internal model that may already have surpassed Astra — but this rests entirely on social media posts cited in Mowshowitz's text and cannot be verified further here.

That leaves several open questions: None of the ARC-AGI-3 numbers have been independently confirmed, and the gap between 99.9 percent on a benchmark and a game-ruining bug in practice deserves an explanation. There is no primary documentation from OpenAI about the model yet, and the details of the weekly usage limit remain unclear until PCMag's full text is available. And if roon's hinted successor exists, Astra's advantages may prove short-lived — unless the gap between benchmark and reality turns out to be the real problem.

Sources: Zvi Mowshowitz, "GPT-6-Astra Can Do Ambitious Things" and PCMag, "I Blew Through My Weekly GPT-6 Limit in an Hour, and It's Still Not AGI".

AIMag.no
AIMag.no
The AIMag.no editorial team covers artificial intelligence, tools, research, and regulation.

Sources

  1. GPT-6-Astra Can Do Ambitious Things - by Zvi Mowshowitzthezvi.substack.com
  2. I Blew Through My Weekly GPT-6 Limit in an Hour, and It's Still Not AGI | PCMagwww.pcmag.com