Anthropic Claims Sonnet 5.5 Beats Its Own Flagship on a Coding Benchmark — at Half the Opus Price
On Monday, 28 September 2026, Anthropic declared Claude Sonnet 5.5, the second model in the Claude 5.5 family, ready for users — just six days after flagship Opus 5.5.
Anthropic declared on Monday, 28 September 2026 that Claude Sonnet 5.5, the second model in the Claude 5.5 family, is ready for users — just six days after the flagship Opus 5.5. The company describes the model as its most capable mid-tier model ever: over 30% faster than its predecessor, with significantly lower token consumption per task — and by Anthropic's own figures, it even beats Opus 5.5 on certain agentic coding benchmarks, at half the price. The model is available from day one in Anthropic's API, on Amazon Web Services, Google Cloud and Microsoft Azure, according to CNBC's coverage as reported by Ecosistema Startup.
What's new — and what is merely claimed
The most important thing to keep in mind: all performance figures and efficiency claims in this story originate from Anthropic itself, relayed through secondary media. No primary source from the company — model card, system card or pricing page — exists in the public source material yet, and none of the benchmark figures have been independently verified.
With that caveat, the picture is this: Anthropic claims Sonnet 5.5 is around 30% faster than its predecessor Sonnet 5 and uses markedly fewer tokens to do the same work — which, according to the company, makes it roughly 30% cheaper per task, even though the list price is unchanged. Both SiliconANGLE and TechCrunch relay the claims from launch day.
On Anthropic's benchmark numbers, Sonnet 5.5 scored 70.6% on Terminal-Bench 4.0, an agentic coding evaluation, versus its predecessor's 10.3%. Via Ecosistema Startup, which cites Decrypt, it also emerges that Opus 5.5 reportedly achieved 66.4% on the same benchmark — meaning the mid-tier model supposedly beats the company's own flagship. This figure, however, is documented only through a single source chain and lacks corroboration. On GDPval-AA, a test of real-world occupational work, Sonnet 5.5 sits two points below Opus 5.5, according to SiliconANGLE. Both pictures can be true — they concern different benchmarks — but together they draw a more nuanced relationship between flagship and mid-tier than "Sonnet beats Opus" alone.
Price and economics: same list price, claimed lower cost
Sonnet 5.5 is priced like its predecessor: $2 per million input tokens, $10 per million output tokens and $0.20 per million cache reads, confirmed by SiliconANGLE. By comparison, Opus 5.5, launched on 22 September, costs $4 per million input tokens and $20 per million output tokens — double.
This is what makes the token-efficiency claim economically significant: if a model solves a task with roughly 30% fewer tokens, the actual cost per job falls correspondingly, without any list price changing. But this is Anthropic's own calculation, not an independent measurement, and no external figures yet confirm the token consumption in practice.
Safety guardrails: the first Sonnet with "large model" measures
One of the most consequential changes for developers is barely mentioned in the marketing: Sonnet 5.5 is the first Sonnet model to launch with cybersecurity measures and guardrails on par with the more capable models. TechCrunch relays Anthropic's claim that the model has "comparable" cyber capabilities to Opus 5, and that it is therefore the first Sonnet covered by the same safety measures that apply to Fable and Opus — SiliconANGLE describes the measures as developed for "Mythos-class models and Fable."
In practice, this means the guardrails can trigger when queries touch cybersecurity, biology or other sensitive content. Anthropic says most software development and bio work tasks are unaffected. One outlet, Ecosistema Startup, claims that sensitive cyber questions are automatically redirected to Opus 4.8 — but this claim is documented only at that single outlet and should be treated as unverified until more emerges.
For teams building agentic systems in security testing, biotechnology or adjacent fields, the message is that the Sonnet class, starting with this model, is no longer an "easy" variant free of such restrictions. It is a change to reckon with from day one.
Watermarking with the EU AI Act in mind
Sonnet 5.5 also ships with invisible text watermarking. When Anthropic launched the watermarking technology, the company stated that it is not meant to affect text quality or readability, but is intended to increase the likelihood that AI-generated text can be detected. The reason, according to the company, is compliance with global regulations, including the EU AI Act.
How the watermark works in practice, how robust it is, and whether the writing quality is genuinely unaffected rest solely on Anthropic's own statements. For now, no independent analysis of the technology exists in the source material. For European organizations, though, the point is concrete: one of the most widely used mid-tier models now delivers text with built-in traceability as standard.
Availability and the context around the launch
The model is available from launch day in Anthropic's own API and at the three major cloud platforms — AWS, Google Cloud and Microsoft Azure — according to CNBC's coverage as reported by Ecosistema Startup. That lowers the barrier for organizations already running Claude workloads that only need to swap a model identifier.
The context is worth noting. Reuters writes that this is the second model in the Claude 5.5 family, launched while Anthropic is building toward a planned IPO, and that enterprise customers account for roughly 80% of the company's business — a figure attributed to Anthropic itself. A cheaper, faster mid-tier model is precisely the kind of product that hits this customer segment. Reuters also describes Sonnet 5.5 as a faster, cheaper complement to Opus 5.5, and states that Haiku 5.5 is coming "in the coming weeks" — with no firm date.
The launch also comes in the middle of a price-and-performance race against OpenAI, which puts the measurements in perspective: in a market where customers choose between models every week, every percentage point on a benchmark and every dollar per task counts.
The early customer picture — with caveats
Zendesk, which tested the model ahead of launch, offers a positive but interested assessment. "We fed Claude Sonnet 5.5 hundreds of real support cases, both responses and escalation requests. It made fewer misjudgments and resolved tickets faster than the Claude models we use in production today. Tickets were handled 20% faster," said Abhinay Kathuria, director of AI at Zendesk, according to SiliconANGLE.
This is an early-adopter test from a customer with its own interest in appearing to be an efficient AI user, and the figures are not independently verified. But they point to the use case Anthropic itself emphasizes: everyday, high-volume tasks where speed and token economics matter more than maximum reasoning power.
What remains to be answered
Three questions remain open after the launch. First, all performance figures lack independent verification — the 70.6% on Terminal-Bench 4.0, the 30% claims on speed and cost, and the relationship between Sonnet 5.5 and Opus 5.5, where the Opus figure on Terminal-Bench is documented only through a single source chain. Second, the watermark's real detectability and any quality effects are unknown outside Anthropic's own promises. Third, the market awaits Haiku 5.5, which at launch was dated only to "the coming weeks."
For developers choosing a model this week, the practical conclusion is nonetheless clear: Sonnet 5.5 offers Opus-adjacent performance on select coding benchmarks — at least according to the vendor — at half the price, but with stricter safety guardrails and standard watermarking than any previous Sonnet.

