Back
AI News

Grok 4.5 Lost Four of Five Evals. xAI Called It Its Strongest Model Ever.

xAI's own launch page shows Fable beating Grok 4.5 on four of five software evals — 70 vs. 53 on DeepSWE 1.1. Yet the model shipped as xAI's strongest ever, and the market didn't punish it. An analysis of what 'trained alongside Cursor' actually buys.

AIMag.no
AIMag.no
August 31, 2026 · 6 min
Paper cutout bar chart with three bars; the shortest, red bar has a string lifting it upward out of frame, while the tall black winning bar sits untouched.

Two things sit side by side on xAI's own launch page. One is the headline: Grok 4.5 is "our strongest model ever." The other is the benchmark charts, where Fable beats Grok 4.5 on four of the five listed software evals — including DeepSWE 1.1, where Fable scores 70 percent to Grok 4.5's 53 (xAI, primary source).

This is likely the first frontier launch that doesn't pretend to top the table. And the market has been careful: J.P. Morgan has reportedly become "increasingly positive" on SpaceX's Grok bet, per Seeking Alpha, and MSN reports that SpaceX stock is getting a "Grok-fueled boost," citing the company's large revenue targets for its AI business.

The numbers xAI itself published

The model launched, per xAI's own news page dated July 16, 2026 (primary source), as the flagship for coding, agentic tasks, and knowledge work. The figures on that same page tell a split story. On DeepSWE 1.0, Grok 4.5 scores 62.0 percent, against Fable (max) at 66.1 and GPT 5.5 (xhigh) at 64.31. On DeepSWE 1.1, it's 53 against Fable's 70 and GPT 5.5's 67. On SWE Bench Pro, Grok 4.5 solves 64.7 percent, against Fable's 80.4. On Terminal Bench 2.1, Grok 4.5 sits at 83.3 — a tenth of a percentage point behind GPT 5.5 and one behind Fable.

Then the one eval where Grok 4.5 wins: SWE Marathon, with a 29.0 percent discovery rate (pass@1) against Opus 4.8's 26.0 and Fable's 24.0.

Worth noting: the competitor figures are not xAI's own measurements. They are drawn, xAI writes, from the respective developers' published system cards and leaderboards, and the evals were partly created and run by Datacurve using each model provider's own harnesses. In other words: vendor-selected frameworks, not neutral measures. But publishing them was xAI's own call — and the company chose to publish.

Why the numbers look this way

xAI's description of the training (primary source) explains much of the profile. Grok 4.5 was trained on tens of thousands of NVIDIA GB300 GPUs, with reinforcement learning over hundreds of thousands of tasks centered on multi-step software work. The stack is built for highly asynchronous training: agentic rollouts can run for many hours while learning continues across tens of thousands of GPUs. The goal, per xAI, is "per-token intelligence" — smarter, more efficient reasoning on real engineering tasks.

There's a logic here: a model optimized for long, multi-step agent runs can lose the short, scored evals and still do the job better when tasks run for hours. The SWE Marathon result, where Grok 4.5 leads, points the same direction.

But the most concrete strategic detail is a different one: Grok 4.5 was "trained alongside Cursor," according to xAI, which links to Cursor's own blog about the collaboration. In practice, that means the model is calibrated against the workflow developers are already sitting in — and the distribution channel ships with the model. This is a product strategy, not a benchmark strategy. The model's competitive position is "good enough, and present where the work happens," not "number one on the board."

The market priced a platform, not a table position

The consequence is well documented in the financial coverage. J.P. Morgan says it has become "increasingly positive" on SpaceX's Grok bet (Seeking Alpha). MSN describes the stock as "Grok-fueled" and points to the company's revenue ambitions for AI. Memeburn frames the same pattern from the product side: Grok Voice 2.0, Grok 4.5, and voice controls in Tesla together turn Grok from a chatbot into a broader AI platform.

In other words, the investment story isn't about whether Grok 4.5 beats Fable. It's about Grok as a revenue platform with multiple distribution surfaces. The day after the 4.5 launch, xAI followed up with Grok 4.6, per Yourstory and Freepress Journal, with a focus on agentic tasks and coding — a cadence that confirms the strategy is rhythm, not crown. One nuance: several sources use the name "SpaceXAI" for the company, a fusion of SpaceX and xAI that reflects how the two are woven together in the market narrative; the sources themselves are inconsistent about the company name and stock ticker.

The counterargument: distribution beats benchmarks — for now

There is a real risk in reading this as victory. If benchmark leadership actually predicts developers' default model choice, "good enough plus distribution" won't hold against Fable and GPT 5.5. A single-digit percentage-point gap can be noise in vendor-selected harnesses; a 17-point gap, as on DeepSWE 1.1 (53 vs. 70), is not noise — and Memeburn's Grok 4.6 headline asks the question openly: "But Is It Really Number One?". The company's own numbers alone give no basis for claiming Grok leads on anything but one of five evals.

But the audit side of the ledger deserves equal weight: xAI published the numbers showing its model losing. That is a rare move in an industry where benchmark charts are otherwise curated to win. And the market rewarded the openness with attention, not punishment — though none of the available sources document a specific share-price move tied to launch day; what we know is that J.P. Morgan turned more positive and that MSN reports a "Grok-fueled boost."

That leaves the industry with an unanswered question: What does "best model" mean when a third-place model can be a winning launch — and what happens to the value of benchmark leaderboards when companies no longer need to sit at the top of them?


Visual direction: A printed leaderboard chart run off on stock ticker tape, with the Grok 4.5 row sitting below Fable, photographed in an editorial still-life style.

Hero image prompt: Editorial still life photograph: a long printed benchmark bar chart on matte uncoated paper, showing a third-ranked bar clearly below a leading bar, printed on continuous tractor-feed stock ticker paper curling off the edge of a dark steel desk, single hard directional daylight from the left, generous negative space upper right, restrained palette of warm paper, black ink and one signal-red accent bar, subtle film grain, physically plausible shadows, no screens, no robots, no neon, no visible readable text beyond abstract bars.

Caption: xAI itself published benchmark charts showing Grok 4.5 below Fable on four of five evals — and still called it its strongest model ever.

Alt text: A printed benchmark chart on ticker tape showing a third-ranked bar clearly below a leading red bar, photographed in hard daylight against a dark steel desk.

AIMag.no
AIMag.no
The AIMag.no editorial team covers artificial intelligence, tools, research, and regulation.