Claude Haiku 5.5: 90% Cheaper, Effort Levels, and the Best Small Model for Subagents
Anthropic released Claude Haiku 5.5 on October 7, 2026: $0.10 input and $0.50 output per million tokens, a 1M-token context window, adjustable effort for the first time on a Haiku model, and benchmark scores that beat GPT-6 Luna. Here's the pricing math, the benchmarks, working code, and how to use it as a subagent under Opus 5.5 and Sonnet 5.5.
- Published
- Reading time
- 10 min read

On this page
TL;DR: On October 7, 2026, Anthropic released Claude Haiku 5.5 (claude-haiku-5-5), which it calls "the cheapest, fastest, and most capable small model we've ever released." It costs $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100K tokens, which is 90% cheaper per token than Haiku 4.5 and about 75% cheaper on a real workload. It has a 1M-token context window, 128K max output, a June 2026 knowledge cutoff, and it's the first Haiku model with adjustable effort (low to max). On Anthropic's published benchmarks it beats OpenAI's GPT-6 Luna at the same price, and Anthropic is positioning it as the subagent that runs next to Opus 5.5 and Sonnet 5.5. Below: pricing math, benchmarks, working code, a subagent pattern, and a migration checklist.

What Anthropic Announced
The launch went out on Anthropic's blog and across social. Daniela Amodei, Anthropic's co-founder and President, posted it on LinkedIn with the short version: Haiku 5.5 is about 75% cheaper than Haiku 4.5 and much more capable, it's "exceptional" at repetitive tasks, and it "makes a great subagent, too." With it, the full Claude 5.5 family of Haiku, Sonnet and Opus is out, and Anthropic says all three meet or exceed its earlier models on alignment, honesty and safety testing.
Four things shipped on the same day:
- 01Claude Haiku 5.5, on the Claude Platform, Amazon Bedrock, Google Cloud and Microsoft Foundry
- 02A 50% price cut on Sonnet 5.5 cache reads, from $0.20 to $0.10 per million tokens
- 03Monthly API credits for Claude Max and Team subscribers
- 04Beta computer use and browser use support in the Python and TypeScript SDKs
Claude Haiku 5.5 Specs at a Glance
| Spec | Claude Haiku 5.5 |
|---|---|
| API model ID | claude-haiku-5-5 |
| Amazon Bedrock ID | anthropic.claude-haiku-5-5 |
| Release date | October 7, 2026 |
| Context window | 1M tokens |
| Max output | 128K tokens (300K on the Batch API with the output-300k-2026-03-24 beta header) |
| Knowledge cutoff | June 2026 |
| Input / output | Text and image in, text out |
| Thinking | Adaptive, on by default |
| Effort levels | low, medium (default), high, xhigh, max |
| Retirement | Not before October 7, 2027 |
Claude Haiku 5.5 Pricing
This is the headline. Pricing is tiered by prompt length: requests up to 100K input tokens get the low rate, and longer ones cost 5x more but are still half the price of Haiku 4.5.
| Per million tokens | Haiku 5.5 (up to 100K) | Haiku 5.5 (over 100K) | Haiku 4.5 | Sonnet 5.5 |
|---|---|---|---|---|
| Input | $0.10 | $0.50 | $1.00 | $2.00 |
| Output | $0.50 | $2.50 | $5.00 | $10.00 |
| Cache writes | $0.125 | $0.625 | $1.25 | $2.50 |
| Cache reads | $0.01 | $0.05 | $0.10 | $0.10 (was $0.20) |
The Batch API takes another 50% off these rates.
90% cheaper vs. 75% cheaper: which number is right?
Both. Per token, Haiku 5.5 is 90% cheaper than Haiku 4.5 up to 100K tokens and 50% cheaper above that. But Haiku 5.5 uses Anthropic's updated tokenizer, which produces somewhat more tokens for the same text, so Anthropic's average workload estimate is about 75% cheaper. According to VentureBeat, about 90% of Haiku 4.5 requests already fit under the 100K threshold, so most teams will see the lower tier.
A worked cost example
Say a classification pipeline runs 1 million requests a month, each with a 3,000-token prompt and a 500-token answer:
| Input cost | Output cost | Monthly total | |
|---|---|---|---|
| Haiku 4.5 | 3B tokens x $1.00 = $3,000 | 0.5B x $5.00 = $2,500 | $5,500 |
| Haiku 5.5 | 3B tokens x $0.10 = $300 | 0.5B x $0.50 = $250 | $550 (before tokenizer overhead) |
Even if the new tokenizer adds a noticeable share of extra tokens, you're still at a fraction of the old bill. Add prompt caching on a shared system prompt (cache reads at $0.01 per million tokens) and the input side nearly disappears.
How it compares on price
| Model | Input / output per MTok |
|---|---|
| Claude Haiku 5.5 | $0.10 / $0.50 (up to 100K tokens) |
| GPT-6 Luna | $0.10 / $0.50 (higher above 272K) |
| Gemini 3.5 Flash-Lite | $0.30 / $2.50 |
| Gemini 3.8 Flash | $0.75 / $3.75 (promo through Dec 31, 2026) |
| Grok 4.3 | $1.25 / $2.50 |
| Claude Sonnet 5.5 | $2.00 / $10.00 |
Prices from VentureBeat's launch coverage. They exclude caching, batch discounts and negotiated rates.
Haiku 5.5 matches GPT-6 Luna exactly on list price. So the real question is capability.
Claude Haiku 5.5 Benchmarks
Anthropic's published numbers, with Haiku 4.5, OpenAI's GPT-6 Luna, and Sonnet 5.5 for reference:
| Benchmark | Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | Sonnet 5.5 |
|---|---|---|---|---|
| GDPval-AA v2.1 (Elo, knowledge work) | 1620 | 735 | 1437 | 1840 |
| AA-Briefcase v1.1 | 1578 | 614 | 1336 | 1824 |
| OSWorld 2.1 (computer use, offline subset) | 72.4% | 15.7% | 48.9% | 83.9% |
| Humanity's Last Exam (no tools) | 45.9% | 10.2% | n/a | 56.9% |
| Humanity's Last Exam (with tools) | 57.4% | 18.7% | n/a | 64.5% |
| Terminal-Bench 4.0 (agentic coding) | 39.2% | 0.0% | 16.4% | 70.6% |
| FrontierCode 1.1 (Main) | 46.4% | n/a | 42.4% | 52.1% |
| Chartography (no tools) | 46.4% | 6.4% | 29.1% | 61.6% |
What stands out:
- It beats GPT-6 Luna on every benchmark where both have a score, at the same price. The gaps are largest on computer use (72.4% vs 48.9% on OSWorld) and agentic coding (39.2% vs 16.4% on Terminal-Bench 4.0).
- The jump from Haiku 4.5 is huge. OSWorld goes from 15.7% to 72.4%, and Terminal-Bench goes from 0% to 39.2%. That's not a refresh, it's a different class of model.
- Sonnet 5.5 still leads on hard agentic coding. 70.6% vs 39.2% on Terminal-Bench is a real gap. Anthropic says so directly: use Sonnet 5.5 or Opus 5.5 for complex agentic coding, and Haiku 5.5 for narrower tasks.
- Effort matters a lot. According to VentureBeat, Haiku's Terminal-Bench score is about 39% at maximum effort and about 20% at the default
medium. Benchmark headlines usually reflect the top setting, so check which level your evals use.
These are vendor-reported numbers. Run your own evals before switching a production workload.
Effort Levels: The Biggest Change for Developers
Haiku 5.5 is the first Haiku model with the effort parameter. Effort controls how many tokens the model spends, including thinking, text and tool calls, so it's now your main control for quality, latency and cost on a single model.
| Effort | When to use it on Haiku 5.5 |
|---|---|
low | Chat, short tool calls, simple high-volume requests (classification, routing, extraction). Cheapest and fastest. |
medium | The default. Most work, including agentic coding. |
high | Knowledge work, longer agent tasks, strict instruction following. |
xhigh / max | Only where your evals show a gain. At this point, compare against Sonnet 5.5 on cost and speed. |
One thing to watch: Anthropic's docs say that at low, in long agent prompts, the model is more likely to skip a search, stop early, or skip a check. Use low for short, well-defined jobs, not long autonomous loops.
Calling Claude Haiku 5.5 with the effort parameter (TypeScript)
import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic();
const response = await client.messages.create({
model: "claude-haiku-5-5",
max_tokens: 1024,
output_config: { effort: "low" },
messages: [
{
role: "user",
content: "Classify this support ticket as billing, bug, feature_request or other: 'I was charged twice this month.'",
},
],
});
const text = response.content.find((b) => b.type === "text");
console.log(text?.text);Python
import anthropic
client = anthropic.Anthropic()
response = client.messages.create(
model="claude-haiku-5-5",
max_tokens=1024,
output_config={"effort": "low"},
messages=[
{"role": "user", "content": "Summarize this changelog in three bullet points: ..."}
],
)
for block in response.content:
if block.type == "text":
print(block.text)Turning thinking off
Thinking is on by default and counts toward max_tokens, so leave room for it. For the fastest possible responses you can send thinking: { type: "disabled" }, but only at high effort or below. At xhigh or max it returns a 400 error. Also, with thinking disabled you can't change effort mid-conversation.
const fast = await client.messages.create({
model: "claude-haiku-5-5",
max_tokens: 256,
thinking: { type: "disabled" },
output_config: { effort: "low" },
messages: [{ role: "user", content: "Extract the order ID from: 'Order #A-99812 never arrived.'" }],
});Using Claude Haiku 5.5 as a Subagent
This is where Haiku 5.5 is most useful. The pattern: a bigger model (Opus 5.5 or Sonnet 5.5) plans and makes decisions, and it hands the repetitive, parallel work to many cheap, fast Haiku 5.5 calls.
Anthropic's launch customers describe exactly this:
- Rogo uses a Haiku 5.5 subagent to pull segment revenue from 10-K filings while a larger model builds the deck.
- Cognition uses Haiku 5.5 as the sidekick in Devin Fusion. With Opus 5.5 leading, it reports a FrontierCode score of 66.2 at lower cost and latency.
Here's a minimal orchestrator / subagent setup in TypeScript. Opus 5.5 splits the job, Haiku 5.5 handles each piece in parallel, and Opus 5.5 combines the results:
import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic();
const textOf = (msg: Anthropic.Message) =>
msg.content.flatMap((b) => (b.type === "text" ? [b.text] : [])).join("");
// 1. Haiku 5.5 subagent: one narrow, repeatable job per call.
async function extractFacts(document: string): Promise<string> {
const msg = await client.messages.create({
model: "claude-haiku-5-5",
max_tokens: 2048,
output_config: { effort: "low" },
system: "Extract revenue, growth and risk factors as terse bullet points. No commentary.",
messages: [{ role: "user", content: document }],
});
return textOf(msg);
}
// 2. Opus 5.5 orchestrator: reasons over the subagents' output.
export async function analyze(documents: string[], question: string) {
const facts = await Promise.all(documents.map(extractFacts));
const msg = await client.messages.create({
model: "claude-opus-5-5",
max_tokens: 8192,
messages: [
{
role: "user",
content: `Question: ${question}\n\nExtracted facts:\n${facts.join("\n---\n")}`,
},
],
});
return textOf(msg);
}Why this works: each extraction prompt stays well under the 100K tier, every call shares a cacheable system prompt, and the expensive model only sees condensed facts instead of every raw document. In production, add a concurrency limit and retries around Promise.all. I covered those patterns in building agentic AI systems and agentic error recovery and observability.
Haiku 5.5 vs Sonnet 5.5 vs Opus 5.5: Which Should You Use?
| Claude Haiku 5.5 | Claude Sonnet 5.5 | Claude Opus 5.5 | |
|---|---|---|---|
| Best for | High-volume, latency-sensitive work: classification, extraction, routing, summaries, compaction, live support, browser use | The best balance of speed and intelligence | Long-running agentic coding and knowledge work |
| Price (in / out per MTok) | From $0.10 / $0.50 | $2 / $10 | $4 / $20 |
| Default effort | medium | high | medium |
| Speed | Fastest | Fast | Moderate |
| Context | 1M | 1M | 1M |
A simple rule: if the task is narrow and you run it many times, start with Haiku 5.5. If it needs long, autonomous, multi-step reasoning, use Sonnet 5.5 or Opus 5.5, and let Haiku be its subagent.
Anthropic notes that Haiku 5.5 is its fastest model at standard speed. Opus models in Fast Mode are still faster.
What Launch Customers Reported
These come from Anthropic's announcement:
- Asana: over 30% lower latency on task completions and up to 2.5x faster inference per agent turn.
- HubSpot: 92.8% average on its CRM evaluation suite over three runs, the best it has seen from a small model.
- AlphaSense: 0.84 vs 0.76 for Haiku 4.5 across 400 queries, a statistically significant gain.
- Box: 11 points higher than Haiku 4.5 at about half the latency.
Migrating from Claude Haiku 4.5 to 5.5
The model ID swap is one line, but a few behaviors changed. Here's the checklist I'd run:
- 01Change the model ID from
claude-haiku-4-5toclaude-haiku-5-5(Bedrock:anthropic.claude-haiku-5-5). - 02Set effort explicitly. The default is
medium. For classification and extraction pipelines, testlowfirst because it's cheaper and faster. - 03Raise
max_tokens. Thinking is on by default and counts toward it. A limit tuned for Haiku 4.5 can cut off answers. - 04Re-measure token counts. The updated tokenizer uses more tokens for the same text, so update cost dashboards and any truncation logic that counts tokens.
- 05Keep prompts under 100K wherever you can. Crossing it multiplies the per-token price by 5x.
- 06Cache shared prefixes. Cache reads at $0.01 per million tokens make long, reused system prompts almost free.
- 07Check cybersecurity prompts. Haiku 5.5's cyber safeguards are stricter than Haiku 4.5's (but looser than Sonnet 5.5's). It allows more defensive work but still blocks penetration testing and attacker-style techniques. Teams with legitimate needs can apply to Anthropic's Cyber Verification Program or Life Sciences program.
- 08Run your evals at more than one effort level before you cut over. Anthropic's own numbers show big differences between
mediumandmax.
The Other Launch-Day Changes
Sonnet 5.5 cache reads cut in half
Sonnet 5.5 cache reads dropped from $0.20 to $0.10 per million tokens. Agent loops resend the same long context every turn, so Anthropic estimates this makes Sonnet 5.5 about 20% cheaper on most agentic workloads. You don't need to change any code to get it. According to VentureBeat, some Google Cloud and Azure customers get the cut a few days later.
Monthly API credits for Max and Team
Claude subscriptions now include API credits that work on any model through Anthropic's platform:
| Plan | Monthly API credit |
|---|---|
| Max 5x | $100 |
| Max 20x | $200 |
| Team | Up to $500, pooled |
If you pay for Max and also prototype against the API, that covers a lot of Haiku 5.5 calls: at $0.10 per million input tokens, $100 buys about a billion input tokens.
Computer use and browser use in the SDKs
The Python and TypeScript SDKs added beta support for computer use and browser use. Together with Haiku 5.5's 72.4% on OSWorld 2.1, that makes cheap browser automation agents practical.
My Take
Small models have usually forced a trade: cheap enough to run on everything, but not smart enough to trust. Haiku 5.5 removes most of that trade for narrow tasks. When a small model scores 72% on a computer-use benchmark and costs $0.10 per million input tokens, a lot of pipelines that used a mid-size model "just to be safe" can move down a tier.
In AI commerce work like mine at Modelia, that's the high-volume layer: tagging product attributes, routing requests, summarizing catalog data, and checking structured output. That work runs thousands of times a day, and each call is simple. It's a good fit for Haiku 5.5 at low effort, with Sonnet 5.5 or Opus 5.5 kept for the steps that need real reasoning.
The part I'd be careful with is effort. The benchmark headlines come from high effort settings, while the default is medium, and low can stop early on long agent tasks. Pick the effort level per task, measure it, and don't assume the launch numbers will carry over to your settings.
If you're building on the Claude API, my guides on Claude agentic workflows in production and the Model Context Protocol pair well with this one. For more on the rest of Anthropic's 2026 lineup, see the Claude Opus 4.8 breakdown and the Claude Fable 5 relaunch.
Sources: Anthropic: Claude Haiku 5.5, Claude models overview, Claude effort docs, VentureBeat launch coverage, 9to5Mac, and Daniela Amodei's LinkedIn post. Header image: public domain (CC0) via Wikimedia Commons.
Harsh Rastogi is an AI Product Engineer at Modelia, building production generative-AI systems for fashion commerce, and the creator of carcode. He writes about AI systems, developer tooling and production engineering at harshrastogi.tech.
Frequently asked questions
What is Claude Haiku 5.5?
Claude Haiku 5.5 is Anthropic's smallest and fastest model in the Claude 5.5 family, released on October 7, 2026. Anthropic calls it the cheapest, fastest and most capable small model it has released. It is built for high-volume, latency-sensitive work such as classification, extraction, routing, summaries and subagent tasks.
How much does Claude Haiku 5.5 cost?
For prompts up to 100K tokens, Claude Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens, with cache reads at $0.01 and cache writes at $0.125. Above 100K tokens it costs $0.50 input and $2.50 output. The Batch API is 50% off.
What is the Claude Haiku 5.5 model ID?
The Claude API model ID is claude-haiku-5-5. On Amazon Bedrock it is anthropic.claude-haiku-5-5, and Google Cloud and Microsoft Foundry use claude-haiku-5-5.
What is the context window of Claude Haiku 5.5?
Claude Haiku 5.5 has a 1M-token context window and up to 128K output tokens on the Messages API. Its reliable knowledge cutoff is June 2026.
Is Claude Haiku 5.5 better than GPT-6 Luna?
On Anthropic's published benchmarks, Haiku 5.5 scores higher than GPT-6 Luna on every benchmark where both have a result, including OSWorld 2.1 (72.4% vs 48.9%), Terminal-Bench 4.0 (39.2% vs 16.4%) and GDPval-AA (1620 vs 1437), at the same $0.10/$0.50 list price. These are vendor-reported numbers, so test on your own workload.
Does Claude Haiku 5.5 support the effort parameter?
Yes. Haiku 5.5 is the first Haiku model with adjustable effort. It supports low, medium, high, xhigh and max, and the default is medium. Set it with output_config.effort on the request.
Should I use Haiku 5.5 or Sonnet 5.5?
Use Haiku 5.5 for narrow, repeated tasks where speed and cost matter, such as classification, extraction, routing and subagents. Use Sonnet 5.5 or Opus 5.5 for complex, long-running agentic coding, where Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0 against Haiku 5.5's 39.2%.
How is Claude Haiku 5.5 75% cheaper if the price per token is 90% lower?
Per token, Haiku 5.5 is 90% cheaper than Haiku 4.5 for prompts up to 100K tokens. Its updated tokenizer produces somewhat more tokens for the same text, so Anthropic estimates the average real workload costs about 75% less.
- Claude
- Anthropic
- Claude Haiku 5.5
- LLM Pricing
- AI Agents
- Developer Tools
Written by Harsh Rastogi, AI Product Engineer and Business AI Head at Modelia. More on AI products, agents and production engineering on LinkedIn.