Skip to main content
Claude Haiku 5.5 AI model, Anthropic's new, more affordable large language model, on a digital screen.

Editorial illustration for Anthropic's Claude Haiku 5.5 model arrives with drastic price cuts

Claude Haiku 5.5: 90% Price Cut on Anthropic AI

Anthropic's Claude Haiku 5.5 model arrives with drastic price cuts

• 4 min read

Anthropic put a new number on the table this week: Claude Haiku 5.5, a small model priced up to 90 percent cheaper than its predecessor for most real-world requests. The company built it for the grunt work of AI deployment, things like summarizing documents, querying databases, sorting support tickets, and running live customer chat, jobs that happen millions of times a day and punish companies for every token spent on a slow or expensive model.

Anthropic says Haiku 5.5 posts real benchmark gains over Haiku 4.5, not just a price cut dressed up as an upgrade, and claims it outperforms OpenAI's budget model GPT-6 Luna on tasks like agentic coding. The model is live now on AWS, Google Cloud, and Azure, and Anthropic paired the release with broader pricing moves: Sonnet 5.5 cache read costs are getting cut in half, and subscribers are getting monthly API credits.

There's a catch buried in the fine print, though. Haiku 5.5 runs on an updated tokenizer, the same kind of change that quietly inflated token counts by about 30 percent when Anthropic rolled out the Opus 4.x line.

Haiku 5.5 is designed for high-volume, cost-sensitive tasks like summarization, database queries, classification, and live customer support, according to Anthropic. On average, the model costs about 75 percent less than Haiku 4.5. For requests with prompts up to 100,000 tokens, which Anthropic says account for roughly 90 percent of all previous Haiku requests, prices drop by up to 90 percent.

Why this matters

For anyone building on Claude, Haiku 5.5 changes the math on what's worth running through a small model versus a flagship one. A 75 percent average price cut, and up to 90 percent on certain token loads, means tasks like classification and customer support queries that teams previously routed to GPT-6 Luna for cost reasons now have a cheaper, apparently more capable Anthropic option. We'd watch two things: whether the benchmark gains over Haiku 4.5 hold up in real production workloads rather than curated tests, and how OpenAI responds on pricing for Luna given the direct comparison Anthropic is inviting.

Availability on AWS, Google Cloud, and Azure from day one matters too, it removes the usual excuse of waiting for infrastructure support before switching providers. For founders running high-volume, low-margin AI features, this is the kind of release that actually moves unit economics, not just a benchmark headline. The bigger story is that neither Anthropic nor OpenAI seems willing to let pricing settle, which is good news for builders and a signal that margins on small models are still being fought over hard.

Common Questions Answered

How much cheaper is Claude Haiku 5.5 compared to its predecessor Haiku 4.5?

Claude Haiku 5.5 costs approximately 75 percent less than Haiku 4.5 on average. For requests with prompts up to 100,000 tokens, which account for roughly 90 percent of all previous Haiku requests, prices drop by up to 90 percent, making it significantly more cost-effective for high-volume deployments.

What specific use cases is Claude Haiku 5.5 designed for according to Anthropic?

Haiku 5.5 is specifically designed for high-volume, cost-sensitive tasks including document summarization, database queries, classification, and live customer support. These are jobs that happen millions of times a day and require models that minimize token spending to reduce operational costs.

How does the pricing structure of Haiku 5.5 affect deployment decisions for AI teams?

The 75 percent average price reduction and up to 90 percent savings on certain token loads change the cost-benefit analysis for routing tasks to small versus flagship models. Teams that previously routed classification and customer support queries to more expensive models like GPT-6 Luna now have a cheaper and apparently more capable Anthropic option available.

What percentage of previous Haiku requests fall within the 100,000 token prompt range where Haiku 5.5 offers maximum savings?

According to Anthropic, approximately 90 percent of all previous Haiku requests have prompts up to 100,000 tokens. This means the vast majority of Haiku users can benefit from the up to 90 percent price reduction that Haiku 5.5 offers for these request sizes.

LIVE23:12ChatGPT's New UI Lets Users Build Custom Tools Like a Plane Seat Explorer