Google's Gemini 4 Argon: 1 million token output, but coding results lag the competition
Google launched its new frontier model Gemini 4 Argon on Wednesday, September 30, 2026, but most paying customers will have to wait: the model goes first to vetted cyber defenders through the company's Fairwind program, with no date for broader… access. The introductory price is $2 per million input tokens and $10 per million output tokens — half the standard price — and the output limit has been raised from 64,000 to 1 million tokens. All benchmark figures come from Google itself, and they are mixed: Argon takes top score on 13 of 19 results in Google's own table, but clearly trails the competition on key coding tests.
What was launched — and by whom
The model was announced in a Google blog post by Koray Kavukcuoglu, senior vice president at Google DeepMind and Google's chief AI architect. Google positions Argon as a frontier model for software development, knowledge work, and cybersecurity (Yahoo Finance).
A backwards rollout: cyber defenders first
The most unusual aspect of the launch is the access model. Argon goes first to "vetted" cyber defenders through the Fairwind program, which Google launched on September 2 for government agencies, Google Cloud customers, and security partners, and which now has more than 650 partners. Google is also participating in the voluntary US government process for pre-release model access (AIStockWire).
Google says the company will continue gathering feedback from early testers and adjusting the safety guardrails before a broader launch. When it comes, it will start with paying API customers and subscribers to Google AI Ultra, the company's top consumer subscription. Google gave no date (MadRobot).
Most strikingly: Google says that trusted defenders and its own teams get a variant «without cyber safety guardrails, so they can leverage its full frontier-level cyber defense capabilities» (AIStockWire). In other words, the most capable version of a frontier model in the cyber domain is reserved for a controlled circle — according to Google, precisely because the capabilities could otherwise be misused. According to Futunn, the monitoring measures mean the model can be terminated if it deviates from its intended scope of activity. Tulsee Doshi, senior director and product lead for Gemini, says: «We have observed that these protective measures are effective» (Futunn).
Pricing and capacity
Google is throwing itself into a price competition. The introductory price is $2 per million input tokens and $10 per million output tokens. After the introductory period, prices double to $4 and $20, and Google has not said how long the introductory period lasts. Cached input is 95% cheaper (Yahoo Finance; MadRobot; AIStockWire). AIStockWire estimates that the price amounts to roughly one-fifth of GPT-6 Astra (AIStockWire).
In addition, Google has raised Argon's output limit from 64,000 to 1 million tokens, described as a response to complex problem-solving scenarios (Stocktwits via TradingView).
The benchmarks: strong on Google's table, weaker head-to-head against competitors
In the benchmark table Google published, Argon holds top score on 13 of 19 results against GPT-6 Astra, Claude Fable 5.1, and Claude Opus 5.5, and ties with GPT-6 Astra on one. All the figures are Google's own and thus not independently verified, and the company has published its evaluation methodology alongside them (AIStockWire).
But the table also contains clear losses. On FrontierSWE v2, Argon scores 55.0%, against Astra's 65.5% — a gap of 10.5 percentage points. On Terminal-bench 4.0, Argon's 57.4% is almost 9 percentage points behind Opus 5.5's 66.4% (Futunn).
A methodological caveat makes the comparisons even harder to read: the conditions were not standardized. On DeepSWE — where Argon leads with 77.9% — Google used its proprietary mini-swe agent framework to compute Argon's score, while the competitors' results were drawn from public leaderboards and the models' own reports (Futunn). All the figures, in other words, deserve careful reading: they are self-reported, and several of the comparisons do not measure the same thing in the same way.
Google's own use cases — the claim of the critical finding
The use cases Google presents come mainly from internally and partner-driven cyber work. The security company Wiz uses Argon for free scanning of critical public infrastructure, and Google claims the model found a critical vulnerability that exposed personal data in healthcare software used by hospitals worldwide — a vulnerability earlier models are said to have missed (MadRobot). This is Google's claim, not independently verified.
Internally, Argon is said by Google to have freed up more than 300 TiB of memory in data centers by analyzing cluster performance telemetry and implementing optimizations, with an estimated potential of 500 TiB to 1 PiB (Futunn). The same applies here: the figures are Google's own, from its own environments, without external control.
Documented skepticism — and Google's response
Simultaneously with the launch, Bloomberg reported, citing people familiar with the matter, that some Google employees believe Gemini 4 performs well on benchmarks but struggles with some real-world coding tasks. Google disputes that characterization (MadRobot). The contradiction is worth holding onto: the company itself highlights coding achievements, while the internal doubts point to exactly the gap between self-reported benchmark figures and practical performance — a gap reinforced by the non-standardized comparison conditions above.
The market response: the sources disagree
Alphabet stock closed Wednesday up 0.9% at $344.15 and traded at $349.45 after hours as of 4:51 p.m. ET, according to AIStockWire (AIStockWire). Stocktwits similarly reports that GOOGL rose around 1% after hours (Stocktwits via TradingView). Yahoo Finance, by contrast, reports that Alphabet stock was flat on the news (Yahoo Finance). The difference may come down to measurement timing — the Yahoo update may have been taken before the after-hours move the other sources capture — but that is unresolved, and there is no unambiguous market read on the launch yet.
What remains open
Three questions remain unanswered. When do paying API customers and AI Ultra subscribers get access — Google has given no date? How long does the introductory price of $2/$10 last before it doubles to $4/$20 — unspecified? And to what extent does the benchmark table actually represent a real lead over the competition, when the figures are Google's own and the comparison conditions partially deviate from what the competitors have reported? Independent benchmarking of Argon will have to come from third parties with access — which the restricted rollout has so far limited.

