Anthropic cuts Haiku 5.5 price to $0.10 per million tokens – but with an important limit
Anthropic launched its new small model Claude Haiku 5.5 on Wednesday at a price matching OpenAI's GPT-6 Luna — but only for requests under 100,000 tokens.
Anthropic launched its new small model Claude Haiku 5.5 on Wednesday at a price matching OpenAI's GPT-6 Luna — but only for requests under 100,000 tokens. Above that threshold, prices rise fivefold, and an independent price calculation suggests long calls can end up at five times Luna's price. The same day, the company halved cache-read prices on Sonnet 5.5 and introduced monthly API credits.
The full price schedule
Haiku 5.5 was launched on 7 October 2026 at $0.10 per million input tokens and $0.50 per million output tokens for requests under 100,000 tokens — exactly the same level as OpenAI's GPT-6 Luna, and a 90 percent drop from its predecessor (VentureBeat).
But the price structure has two tiers. Above 100,000 tokens, input costs $0.50 and output $2.50 per million tokens. That is still a 50 percent reduction from Haiku 4.5, but far from the headline figure. SiliconANGLE reports that Anthropic prices the new model at roughly a quarter of what Haiku 4.5 costs to run (SiliconANGLE).
The threshold is no accident: according to Anthropic, prompts of up to 100,000 tokens accounted for around 90 percent of the old model's requests (SiliconANGLE). Most existing users will therefore land on the cheap tier — but the 10 percent with long contexts, typically agentic jobs that read large volumes of documents, hit the more expensive prices.
What does it actually cost? Working through the difference
The Japanese analysis site XenoSpectrum has calculated what the threshold means in practice. For a single API call with 200,000 input tokens and 20,000 output tokens, Haiku 5.5 costs an estimated $0.15, while GPT-6 Luna costs $0.03 — roughly five times more expensive (XenoSpectrum). The difference arises because Luna's surcharge only kicks in above 272,000 tokens.
It is important to stress that this is a price calculation, not a measured benchmark: the figures follow mechanically from the published rates but assume a specific call format. The point still stands: for workloads with long inputs, "90 percent cheaper" does not transfer — there, Haiku 5.5 is substantially more expensive than the competitor it officially matches.
The 75 percent figure is an estimate with caveats
Anthropic itself states that workloads overall cost roughly 75 percent less than on Haiku 4.5, once request distribution and changes in token consumption are taken into account (VentureBeat). The estimate rests on two assumptions, both from Anthropic itself: that around 90 percent of requests fall below the threshold, and that an updated tokenizer counts the same text as a somewhat higher number of tokens. How much the tokenizer change contributes is not quantified. XenoSpectrum explicitly warns against transferring the 75 percent figure directly to an arbitrary workload.
Performance: stronger numbers, but all from Anthropic
According to Anthropic's own published benchmarks, Haiku 5.5 scores 72.4 percent on OSWorld 2.1 (the offline portion, a test where agents operate a real computer through multi-step tasks) versus 48.9 percent for GPT-6 Luna. On Terminal-Bench 4.0, the figures are 39.2 versus 16.4 percent (SiliconANGLE). The model is the first Haiku with an adjustable effort setting.
The customer Asana reports from pre-release testing that latency on completed tasks was more than 30 percent lower than with the model it uses today, and that inference per agent turn ran up to 2.5 times faster — as stated by Asana staff software engineer Aaron Vinh through Anthropic's launch materials (SiliconANGLE). None of the performance figures have been independently verified; they are repeated by several outlets, but all originate from Anthropic.
TechTarget places the model as best suited for latency-sensitive tasks such as browser use and customer support (TechTarget).
Sonnet 5.5 also gets cheaper
The same day, Anthropic cut cache-read prices on Sonnet 5.5 — launched on 28 September — from $0.20 to $0.10 per million tokens. Since cached tokens make up a large share of what models actually consume, Anthropic expects the cut to remove around 20 percent of the cost of most agentic work on the model (SiliconANGLE). This, too, is the company's own estimate.
Safety, availability and API credits
Anthropic states that Haiku 5.5 showed fewer instances of misaligned behavior than Haiku 4.5. On cybersecurity, the company writes that the model's safeguards "allow a broader range of defensive tasks than our safeguards for Sonnet 5.5, but they still block penetration testing and other techniques more likely to be used by attackers" (Seeking Alpha via MSN).
The model is available on all major platforms, including AWS, Google Cloud and Microsoft Azure. Anthropic is also introducing monthly API credits: $100 per month for Max 5x users, $200 for Max 20x, and up to $500 for Team subscribers (Seeking Alpha via MSN).
The context: a complete family and possible IPO pressure
The launch completes the Claude 5.5 family nine days after Sonnet 5.5, with Opus 5.5 and Sonnet 5.5 released in September. It also comes as Anthropic is expanding its product portfolio ahead of a possible IPO that could come as early as mid-November (Benzinga via MSN). A target valuation of around $2 trillion is also reported — but this is based on sources citing ongoing discussions, is not confirmed by the company, and could change.
Open questions
Several things remain unresolved. It is unspecified how requests of exactly 100,000 tokens are billed, and which tokens determine the threshold — a genuine ambiguity for cost planning. Furthermore, all performance figures were published by Anthropic itself, and the 75 percent savings rest on the company's own data on request distribution and tokenizer changes. And whether the aggressive pricing strategy against GPT-6 Luna is durable or a marketing phase around a possible listing is something none of the sources can say with certainty.

