Grok 4.7: $3.74 per Measured Task, as Measured by Artificial Analysis
SpaceXAI launched the frontier model Grok 4.7 on September 21, 2026, with unchanged unit pricing — $2 per million input tokens and $6 per million output tokens — but independent measurements from Artificial Analysis show the model uses…

Grok 4.7: $3.74 per Measured Task, as Measured by Artificial Analysis
SpaceXAI launched the frontier model Grok 4.7 on September 21, 2026, with unchanged unit pricing — $2 per million input tokens and $6 per million output tokens — but independent measurements from Artificial Analysis show the model uses roughly twice as many tokens per task as its predecessor. Whether the low token price actually translates into cheaper completed tasks remains an open question.
What Was Launched
SpaceXAI announced Grok 4.7 on Monday, September 21, 2026. The model ID is grok-4.7, the context window is 500,000 tokens, and the API price is $2 input / $6 output per million tokens for prompts under 200,000 tokens, according to KintoNavi, which reproduces the company's announcement ([2f46d37c]). The company itself describes the model as its most capable yet for coding and knowledge work (Seeking Alpha, [e782cbc8]).
The model became available the same day in Cursor, Grok Build, and via the API, as well as on OpenRouter, Vercel, and Cloudflare. GitHub Copilot began a phased rollout on September 21 for the Pro, Pro+, Max, Business, and Enterprise subscriptions, with access from the model picker in VS Code, Visual Studio, Copilot CLI, cloud agents, the Copilot app, JetBrains, Xcode, and Eclipse (わど, [06c0e89f]).
The launch came roughly 40 days after Grok 4.6 (Morimoto, [a593694c]) — a very short interval for a new frontier model.
The Company's Own Numbers
According to SpaceXAI's own benchmark table, the largest gain is on Terminal-Bench 4.0, which rose from 20.3% for Grok 4.6 to 38.0% for Grok 4.7 — an increase of 17.7 percentage points (KintoNavi, [2f46d37c]). In the company's own table, Grok 4.7 leads on EEBench and the Harvey Legal Agent Benchmark.
But even in the company's own numbers, competitors lead on several benchmarks: Fable 5.1 beats Grok 4.7 on CursorBench 4.0 (51.8%) and Terminal-Bench 4.0 (57.9%), GPT-5.6 Sol leads DeepSWE v1.1 with 72.7%, and Fable 5.1 leads HealthBench Professional with 62.1% (KintoNavi, [2f46d37c]).
The Independent Measurements
The independent testing lab Artificial Analysis paints a more nuanced picture. On their Intelligence Index, Grok 4.7 at the highest reasoning level, xHigh, scores 46 points — only two more than Grok 4.6. That places SpaceXAI among the four leading AI labs, but the model still trails the top models Claude Fable 5.1 and GPT-6 Astra (heise, [4ebc229b]). At the same time, Artificial Analysis reports that the model comes close to the top on the AA Briefcase and GDPval-AA subtests ([4ebc229b]).
The most striking figure concerns token consumption: At xHigh, Grok 4.7 uses an average of around 81,000 output tokens per task on the Intelligence Index. By comparison, Grok 4.6 xHigh uses around 38,000 tokens, and GPT-6 Astra Max about 27,000. That works out to a cost per task of roughly $3.74 (heise, [4ebc229b]).
This is what makes the launch interesting beyond the usual benchmark tables: the unit price is unchanged, but token consumption per task has roughly doubled. According to Morimoto's analysis, it remains to be seen whether the low unit price actually results in a lower cost for completed tasks ([a593694c]).
An Unexplained Discrepancy in the Numbers
There is also a direct conflict between the company's numbers and the independent ones. SpaceXAI reports 38.0% on Terminal-Bench 4.0, while heise/Artificial Analysis measures 33% — up from 18% for the predecessor. The DeepSWE figures also diverge (71.0% claimed versus 73% measured). The discrepancy is likely due to different reasoning-effort settings or different test setups, but this cannot be confirmed from the available sources, and the discrepancy remains unexplained here.
The Safety Numbers Are Pure Company Claims
SpaceXAI claims that Grok 4.7 is the model it has tested with the strongest refusal and jailbreak resistance, built on an entirely new safety layer. The announcement cites 62.4% on LatchBio's biosafety benchmark and that only 3.3% of risky dual-use prompts get through on HackerBench v0.3, while the refusal rate for legitimate security work is said to be low. The company also offers invite-based red-teaming for selected cybersecurity partners (iClarified, [b5f11e1f]).
These numbers are exclusively the company's own claims and are not independently verified.
Curiosities and Caveats
Some details are worth noting — and treating with caution:
- Grok 4.7 Fast runs the same model on faster infrastructure at double the price for normal context (long context in Cursor is priced separately). The Fast variant is not available via the public xAI API, only in Cursor and Grok Build (わど, [06c0e89f]).
- According to heise, SpaceXAI used anonymized workflow data from Cursor to strengthen the model's coding and agent capabilities, and Cursor is said to have belonged to SpaceX since August ([4ebc229b]). The relationship between SpaceXAI, xAI, SpaceX, and Cursor's ownership is claimed by secondary sources and not confirmed by primary documentation.
- The sourcing situation should be noted: none of the sources in this story are SpaceXAI's own announcement — all numbers trace to secondary reporting that quotes the company, and several of the sources are machine-translated posts. The details about general chat access (grok.com, mobile apps, X) are also thinly sourced.
What You Should Measure Yourself
The practical point for anyone considering switching models is that cost per completed task — not price per token — is the relevant purchasing criterion. The independent measurements suggest that Grok 4.7 solves tasks with far more reasoning text, but favorable traits such as a lower error rate or fewer retry rounds could still make the model cheap in practice. Nobody knows yet.
Before any migration, it makes sense to measure your own workflows: average token consumption per actual task type, first-attempt success rate, latency at relevant reasoning levels, and comparison against both Grok 4.6 and the competitors. Particularly those running large volumes of agentic coding tasks via the API will feel the token consumption directly on the bill — regardless of what the unit price is.
Sources
- SpaceXAI Launches Grok 4.7 for Coding and Knowledge Work - iClarified — www.iclarified.com
- Grok 4.7 Arrives | 'Long-Duration Tasks' and Grok Bot: What to Look for Beyond the 500k Token Limit|わど|AI界隈のマスコット — note.com
- Grok 4.7 Released! When to Use It? A Quick Guide to Its Features and Key Points|Yasuhito Morimoto — note.com
- SpaceXAI's new Grok 4.7 improves on coding, maintains cheaper token rates than peers | Seeking Alpha — seekingalpha.com
- SpaceXAI Releases Grok 4.7: Input Price Remains at $2, Terminal-Bench Increases from 20.3% to 38.0%|KintoNavi|kintone+生成AI — note.com
- Grok 4.7: Günstige API trifft auf hohen Tokenverbrauch | heise online — www.heise.de