Anthropic released Claude Haiku 5.5 on Tuesday, describing it as "the cheapest, fastest, and most capable small model we've ever released." The launch comes with a steep two-tier price cut that drops input costs to $0.10 per million tokens for most requests, a move that puts direct pressure on OpenAI's GPT-6 Luna in the small-model segment and extends a months-long price war across the AI industry.

The announcement, published on Anthropic's website on October 7, 2026, positions Haiku 5.5 as the third model in the Claude 5.5 family, alongside Sonnet 5.5 and Opus 5.5. Reuters characterized the release as Anthropic expanding its model lineup as it prepares for a planned IPO, and the model is already available in GitHub Copilot, Amazon Web Services, Google Cloud, and Microsoft Foundry. The launch is the latest entry in a month of breaking AI news that has seen model releases, price cuts, and policy moves land nearly every day.

A Small Model Built for High-Volume Work

According to Anthropic, Claude Haiku 5.5 is designed for high-volume, cost-sensitive tasks rather than frontier-level reasoning. The company says it "reliably handles quick and repetitive workloads" such as summaries, compaction, database queries, and classification requests.

Anthropic also positions the new model as a companion to its larger systems: Haiku 5.5 "pairs well with Opus 5.5 and Sonnet 5.5 as a subagent on coding work," letting developers route cheap, fast subtasks to Haiku while reserving bigger models for complex reasoning. Because it is Anthropic's fastest model to date at standard speed, the company recommends it for latency-sensitive applications like live customer support and browser use. One footnote adds a caveat: Anthropic's Opus models still run faster in Fast Mode.

Pricing: Up to 90% Cheaper Than Haiku 4.5

The headline change is price. Anthropic says Haiku 5.5 costs around 75% less to run than Haiku 4.5 on average, with a tiered structure that depends on prompt size.

For requests with prompts up to 100,000 tokens — which Anthropic says accounted for roughly 90% of requests to the previous Haiku model — prices are 90% lower than Haiku 4.5. Prompts over 100,000 tokens cost 50% less.

The per-million-token price table from Anthropic's announcement shows the scale of the cut:

  • Input tokens: $0.10 (prompts up to 100K) or $0.50 (over 100K), down from $1.00 flat for Haiku 4.5
  • Output tokens: $0.50 or $2.50, down from $5.00
  • Cache reads: $0.01 or $0.05, down from $0.10
  • Cache writes: $0.125 or $0.625, down from $1.25

At the sub-100K tier, that makes Haiku 5.5's output price one-fifth of Sonnet 5.5's $10.00 per million output tokens.

Anthropic notes one technical detail behind the numbers: Haiku 5.5 uses an updated tokenizer, similar to Sonnet 5.5 and Opus 5.5, which means it uses slightly more tokens per task than Haiku 4.5 did. The company says its average-cost calculation accounts for that difference.

Benchmarks: Ahead of GPT-6 Luna in Most Rows

Anthropic published a benchmark table comparing Haiku 5.5 against Haiku 4.5, GPT-6 Luna, and Sonnet 5.5. On most of the listed evaluations, the new model finishes ahead of Luna and dramatically ahead of its predecessor:

  • GDPval-AA v2.1 (knowledge work): 1620 for Haiku 5.5, versus 735 for Haiku 4.5, 1437 for GPT-6 Luna, and 1840 for Sonnet 5.5
  • AA-Briefcase v1.1: 1578, versus 614 for Haiku 4.5 and 1336 for Luna
  • OSWorld 2.1, offline subset (computer use): 72.4%, versus 15.7% for Haiku 4.5, 48.9% for Luna, and 83.9% for Sonnet 5.5
  • Humanity's Last Exam: 45.9% without tools and 57.4% with tools, versus 10.2% and 18.7% for Haiku 4.5
  • Terminal-Bench 4.0 (agentic coding): 39.2%, versus 16.4% for Luna and 70.6% for Sonnet 5.5
  • FrontierCode 1.1, Main: 46.4%, ahead of Luna's 42.4%

The pattern is consistent: Haiku 5.5 closes much of the gap to GPT-6 Luna and, in computer use, beats it outright, while Anthropic's own Sonnet 5.5 still leads every category. Anthropic's guidance reflects that hierarchy — the company says Sonnet 5.5 and Opus 5.5 "remain better choices for complex agentic coding tasks like those measured by Terminal-Bench 4.0," while Haiku 5.5 targets narrowly scoped work that "might otherwise have been cost-prohibitive with previous versions of Claude."

First Haiku With an Effort Dial

Haiku 5.5 is also the first Haiku-class model to ship with an adjustable effort setting, the cost-versus-intelligence control Anthropic previously offered on its larger models. Users can choose how much compute each request consumes, and Anthropic's charts show benchmark scores at each effort level from Low to Max.

Safety Posture: Tighter Than Haiku 4.5, Looser Than Sonnet

On safety, Anthropic says Haiku 5.5 shows "major improvements across almost all of our alignment evaluations relative to Haiku 4.5," with far fewer instances of misaligned behavior and a lower willingness to cooperate with misuse.

The model's cybersecurity safeguards are more restrictive than Haiku 4.5's, though somewhat less restrictive than those on Anthropic's recent larger models. They permit a wider range of defensive security tasks than Sonnet 5.5's safeguards but still block penetration testing and techniques more likely to be used by attackers. Biology safeguards match those of Sonnet 5, Sonnet 5.5, and Opus 5. Organizations needing broader access can apply to Anthropic's Life Sciences or Cyber Verification Programs. Full details are in the Haiku 5.5 system card.

The Rest of the Announcement: Cheaper Sonnet, Free API Credits

The launch was packaged with several other changes. Cache reads on Claude Sonnet 5.5 now cost $0.10 per million tokens instead of $0.20 — a 50% cut that Anthropic says makes Sonnet 5.5 about 20% cheaper on most agentic work, since cache reads dominate token consumption in agent workflows.

Anthropic is also rolling out monthly API credits this week for subscription customers: $100 per month for Max 5x users, $200 for Max 20x users, and up to $500 pooled across users for Team plans, usable on any model. The credits are aimed at subscribers building tools and agents on the Claude Platform.

Developers get SDK updates as well: Anthropic's Python and TypeScript SDKs now support computer use and browser use in beta, capabilities the company says suit Haiku 5.5's speed-price profile.

A Price War With No End in Sight

The Decoder's coverage framed the launch as proof that "the AI pricing arms race is far from over," and the numbers support that reading. In September, Anthropic and OpenAI launched rival flagship models within hours of each other, complete with competing price cuts. Haiku 5.5 extends that battle to the low end of the market, where small models power the highest-volume, most price-sensitive workloads.

For developers, the immediate effect is straightforward: small-model inference on the Claude platform now starts at $0.10 per million input tokens and $0.50 per million output tokens, with GPT-6 Luna beaten on most of Anthropic's published benchmarks at that price point. Whether OpenAI responds with a Luna price cut — or waits for its next small-model release — will signal how durable this pricing floor really is.

Haiku 5.5 is available now across Anthropic's first-party apps and API, Amazon Web Services, Google Cloud, and Microsoft Foundry, and is already integrated into GitHub Copilot.

Stay Ahead of AI Read more AI news on AI Buzz Wire