Thursday 8 October 2026 556 stories on file Full archive
Daily Edition
newscms

Volume III Edition Daily

Anthropic Unveils Claude Haiku 5.5: A 90% Price Cut on Its Fastest Model

On October 7, 2026, Anthropic released Claude Haiku 5.5, the third model in its Claude 5.5 family. The launch expands Anthropic's lineup with a focus on high-speed, low-cost inference for repetitive work such as…

AI 1,363 words 7 min read

Anthropic Unveils Claude Haiku 5.5: A 90% Price Cut on Its Fastest Model — AI No Image AI
Lead image · Filed 7 October 2026, 23:03

Anthropic Unveils Claude Haiku 5.5: A 90% Price Cut on Its Fastest Model

Introduction

On October 7, 2026, Anthropic released Claude Haiku 5.5, the third model in its Claude 5.5 family. The launch expands Anthropic's lineup with a focus on high-speed, low-cost inference for repetitive work such as classification, summarization, extraction, live support, voice agents, and in-app assistants. The key headline is pricing: for requests under 100,000 tokens, Haiku 5.5 is priced at $0.10 per million input tokens and $0.50 per million output tokens — a 90% reduction compared with Haiku 4.5. For longer prompts above 100,000 tokens, rates fall to $0.50 and $2.50 per million, respectively. Anthropic estimates workloads cost roughly 75% less to run than on Haiku 4.5 once request sizes and tokenization changes are accounted for.

Pricing Structure

The new pricing tiers have a clear boundary at 100,000 tokens. The largest savings apply to the shortest requests, where costs drop by 90%. Above that threshold, savings are 50%. The company also cut Sonnet 5.5 cache-read charges from $0.20 to $0.10 per million tokens, with cache writes and cache reads following the same proportional changes across tiers. Anthropic says about 90% of Haiku 4.5 requests fell below the 100,000-token mark, meaning most existing workloads will see the full 90% reduction.

Charge (per million tokens) Haiku 5.5: <100k Haiku 5.5: >100k Haiku 4.5
Input $0.10 $0.50 $1.00
Output $0.50 $2.50 $5.00
Cache reads $0.01 $0.05 $0.10
Cache writes $0.125 $0.625 $1.25

These rates place Haiku 5.5 in direct competition with other low-cost proprietary models on the short-prompt tier, while the longer-prompt tier remains competitive for tasks that require extended context.

Performance and Benchmarks

Anthropic reports benchmark gains over Haiku 4.5, with particularly notable jumps on computer-use and coding tasks. According to vendor-reported figures, Haiku 5.5 achieves 72.4% on OSWorld 2.1 (offline subset) compared with 15.7% for Haiku 4.5, and 39.2% on Terminal-Bench 4.0 (agent-based coding) at maximum effort (versus 0.0% for its predecessor on that configuration). The model introduces adjustable effort levels, with medium set as the default; results vary by effort setting, with medium scoring around 20% on Terminal-Bench. On knowledge-work benchmarks GDPval-AA v2.1 and AA-Briefcase v1.1, Haiku 5.5 scores roughly double Haiku 4.5. Vendor-reported results should be treated as indicative rather than independently verified, but they reflect substantial improvements in agentic and reasoning tasks.

Use Cases and Positioning

Anthropic positions Haiku 5.5 as a supporting worker model for more capable systems, best suited to high-volume, well-defined tasks where speed and cost dominate. Examples include retrieval for agents, classification, summarization, extraction, live support, voice agents, and in-app assistants. Customer testimonials cited in the launch note improvements in latency and task-completion speed for production workflows, reinforcing its fit for scaled automation rather than open-ended frontier reasoning. The model remains Anthropic's fastest in standard speeds, though larger models can run faster in "fast mode" configurations.

Availability and Additional Updates

Haiku 5.5 launches via Anthropic's platform, AWS, Google Cloud, and Microsoft Azure under the API identifier claude-haiku-5-5. Python and TypeScript SDK updates add beta capabilities for operating browsers and computers. Alongside the Haiku release, Sonnet 5.5 cache-read pricing drops to $0.10 per million tokens. Monthly API credits are also rolling out this week: $100 for Max 5x subscribers, $200 for Max 20x, and up to $500 shared across Team subscriptions (subject to allocation/rollover details in their respective plans). The release also tightens cybersecurity restrictions relative to Haiku 4.5, including blocking penetration testing under standard safeguards; organizations requiring broader access can apply through Anthropic's verification programs.

Conclusion

With Claude Haiku 5.5, Anthropic is targeting the economics of AI at scale. By slashing prices by up to 90% for short prompts while reporting substantial gains on agentic benchmarks, the company is making high-throughput inference far more affordable without abandoning safety guardrails. For teams running thousands of narrow, repetitive tasks each day, the cost-per-request drop could be the difference between experimental and production-scale deployment — a practical shift that aligns capability with affordability.

Author: CMS Publisher Date: October 2026

Market Context: The Race to the Bottom on Price

The Haiku 5.5 launch continues a broader pattern of aggressive price competition among frontier labs. Earlier in the year, Anthropic and OpenAI both cut model prices by roughly half, and Google has positioned Gemini Flash variants as the low-cost tier of its own lineup. In this environment, the short-prompt tier has become the most contested segment: classification, extraction, and summarization jobs are exactly the workloads enterprises run at scale, and a 90% price reduction materially changes what is economical to automate. VentureBeat's coverage of the launch frames the move as Anthropic matching OpenAI's GPT-6 Luna rates at the sub-100k tier while undercutting Google's Gemini Flash-Lite and xAI's Grok models on input pricing.

The comparison is not purely about the sticker rate. Token prices alone do not determine the cost of completing a job — token consumption, retries, and accuracy all affect the final bill. Anthropic's estimate of 75% average workload savings accounts for an updated tokenizer that uses somewhat more tokens for equivalent work, a detail that matters for teams forecasting budgets. Cache pricing follows the same logic: cheaper cache reads and writes reward applications that reuse stored context across requests, which is common in agentic workflows.

Safety and Cybersecurity Guardrails

Anthropic states that Haiku 5.5 is its first Haiku with built-in safeguards for a narrow set of high-risk cybersecurity requests, while most everyday tasks are unaffected. The model tightens restrictions compared with Haiku 4.5, including blocking penetration testing under standard safeguards. Organizations seeking broader cybersecurity or biology access can apply through Anthropic's verification programs. The change signals that even the cheapest tier is not a bare-bones model: the safety layer ships at every price point, an approach consistent with the lab's published stance on scalable oversight.

Developer Perspective

For developers, the practical test is whether Haiku 5.5 can complete a narrowly defined assignment reliably enough to justify repeated use. The lower prices and reported gains make that test more attractive, but the vendor's own benchmark table still supports reserving complex work for larger models. The adjustable effort levels add a useful lever: medium effort by default, with maximum effort reserved for harder tasks. Teams that previously routed everything through Sonnet or Opus can now run a two-tier architecture — Haiku for the long tail of simple calls, larger models for the hard cases — and the pricing makes that architecture viable at volume.

The SDK updates extend the model's reach beyond text generation. Beta capabilities for operating browsers and computers in the Python and TypeScript SDKs mean Haiku 5.5 can serve as the fast layer in computer-use agents, a category that has moved quickly this year. Availability across Anthropic's own platform plus AWS, Google Cloud, and Microsoft Azure means existing cloud commitments can absorb the new model without migration.

Conclusion

With Claude Haiku 5.5, Anthropic is targeting the economics of AI at scale. By slashing prices by up to 90% for short prompts while reporting substantial gains on agentic benchmarks, the company is making high-throughput inference far more affordable without abandoning safety guardrails. For teams running thousands of narrow, repetitive tasks each day, the cost-per-request drop could be the difference between experimental and production-scale deployment — a practical shift that aligns capability with affordability. Readers following AI infrastructure trends can track related coverage in our AI category, where we publish model releases, pricing analysis, and industry reporting.

Images

Jack Clark, co-founder of Anthropic, giving a talk on AI at the Schwarzman Centre in Oxford

Jack Clark, co-founder of Anthropic, speaking at an AI event. The screen behind him shows a chart of model scores over time. Photo via Wikimedia Commons.

Modern workspace with a black keyboard and blue LED lighting

A dark-themed developer workspace with a black keyboard and blue LED accent lighting — the kind of setup where high-volume API calls run. Photo via Wikimedia Commons.

References

Author: CMS Publisher Date: October 2026