Claude Haiku 5.5: 90% Cheaper, Effort Levels, and the Best Small Model for Subagents

    Anthropic released Claude Haiku 5.5 on October 7, 2026: $0.10 input and $0.50 output per million tokens, a 1M-token context window, adjustable effort for the first time on a Haiku model, and benchmark scores that beat GPT-6 Luna. Here's the pricing math, the benchmarks, working code, and how to use it as a subagent under Opus 5.5 and Sonnet 5.5.

    Published
    Reading time
    10 min read
    Claude Haiku 5.5: glowing AI brain over a circuit board, illustrating Anthropic's fastest and cheapest small model
    On this page

    TL;DR: On October 7, 2026, Anthropic released Claude Haiku 5.5 (claude-haiku-5-5), which it calls "the cheapest, fastest, and most capable small model we've ever released." It costs $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100K tokens, which is 90% cheaper per token than Haiku 4.5 and about 75% cheaper on a real workload. It has a 1M-token context window, 128K max output, a June 2026 knowledge cutoff, and it's the first Haiku model with adjustable effort (low to max). On Anthropic's published benchmarks it beats OpenAI's GPT-6 Luna at the same price, and Anthropic is positioning it as the subagent that runs next to Opus 5.5 and Sonnet 5.5. Below: pricing math, benchmarks, working code, a subagent pattern, and a migration checklist.

    Claude Haiku 5.5: Anthropic's fastest and cheapest small model, released October 7, 2026
    Claude Haiku 5.5: Anthropic's fastest and cheapest small model, released October 7, 2026

    What Anthropic Announced

    The launch went out on Anthropic's blog and across social. Daniela Amodei, Anthropic's co-founder and President, posted it on LinkedIn with the short version: Haiku 5.5 is about 75% cheaper than Haiku 4.5 and much more capable, it's "exceptional" at repetitive tasks, and it "makes a great subagent, too." With it, the full Claude 5.5 family of Haiku, Sonnet and Opus is out, and Anthropic says all three meet or exceed its earlier models on alignment, honesty and safety testing.

    Four things shipped on the same day:

    1. 01Claude Haiku 5.5, on the Claude Platform, Amazon Bedrock, Google Cloud and Microsoft Foundry
    2. 02A 50% price cut on Sonnet 5.5 cache reads, from $0.20 to $0.10 per million tokens
    3. 03Monthly API credits for Claude Max and Team subscribers
    4. 04Beta computer use and browser use support in the Python and TypeScript SDKs

    Claude Haiku 5.5 Specs at a Glance

    SpecClaude Haiku 5.5
    API model IDclaude-haiku-5-5
    Amazon Bedrock IDanthropic.claude-haiku-5-5
    Release dateOctober 7, 2026
    Context window1M tokens
    Max output128K tokens (300K on the Batch API with the output-300k-2026-03-24 beta header)
    Knowledge cutoffJune 2026
    Input / outputText and image in, text out
    ThinkingAdaptive, on by default
    Effort levelslow, medium (default), high, xhigh, max
    RetirementNot before October 7, 2027

    Claude Haiku 5.5 Pricing

    This is the headline. Pricing is tiered by prompt length: requests up to 100K input tokens get the low rate, and longer ones cost 5x more but are still half the price of Haiku 4.5.

    Per million tokensHaiku 5.5 (up to 100K)Haiku 5.5 (over 100K)Haiku 4.5Sonnet 5.5
    Input$0.10$0.50$1.00$2.00
    Output$0.50$2.50$5.00$10.00
    Cache writes$0.125$0.625$1.25$2.50
    Cache reads$0.01$0.05$0.10$0.10 (was $0.20)

    The Batch API takes another 50% off these rates.

    90% cheaper vs. 75% cheaper: which number is right?

    Both. Per token, Haiku 5.5 is 90% cheaper than Haiku 4.5 up to 100K tokens and 50% cheaper above that. But Haiku 5.5 uses Anthropic's updated tokenizer, which produces somewhat more tokens for the same text, so Anthropic's average workload estimate is about 75% cheaper. According to VentureBeat, about 90% of Haiku 4.5 requests already fit under the 100K threshold, so most teams will see the lower tier.

    A worked cost example

    Say a classification pipeline runs 1 million requests a month, each with a 3,000-token prompt and a 500-token answer:

    Input costOutput costMonthly total
    Haiku 4.53B tokens x $1.00 = $3,0000.5B x $5.00 = $2,500$5,500
    Haiku 5.53B tokens x $0.10 = $3000.5B x $0.50 = $250$550 (before tokenizer overhead)

    Even if the new tokenizer adds a noticeable share of extra tokens, you're still at a fraction of the old bill. Add prompt caching on a shared system prompt (cache reads at $0.01 per million tokens) and the input side nearly disappears.

    How it compares on price

    ModelInput / output per MTok
    Claude Haiku 5.5$0.10 / $0.50 (up to 100K tokens)
    GPT-6 Luna$0.10 / $0.50 (higher above 272K)
    Gemini 3.5 Flash-Lite$0.30 / $2.50
    Gemini 3.8 Flash$0.75 / $3.75 (promo through Dec 31, 2026)
    Grok 4.3$1.25 / $2.50
    Claude Sonnet 5.5$2.00 / $10.00

    Prices from VentureBeat's launch coverage. They exclude caching, batch discounts and negotiated rates.

    Haiku 5.5 matches GPT-6 Luna exactly on list price. So the real question is capability.

    Claude Haiku 5.5 Benchmarks

    Anthropic's published numbers, with Haiku 4.5, OpenAI's GPT-6 Luna, and Sonnet 5.5 for reference:

    BenchmarkHaiku 5.5Haiku 4.5GPT-6 LunaSonnet 5.5
    GDPval-AA v2.1 (Elo, knowledge work)162073514371840
    AA-Briefcase v1.1157861413361824
    OSWorld 2.1 (computer use, offline subset)72.4%15.7%48.9%83.9%
    Humanity's Last Exam (no tools)45.9%10.2%n/a56.9%
    Humanity's Last Exam (with tools)57.4%18.7%n/a64.5%
    Terminal-Bench 4.0 (agentic coding)39.2%0.0%16.4%70.6%
    FrontierCode 1.1 (Main)46.4%n/a42.4%52.1%
    Chartography (no tools)46.4%6.4%29.1%61.6%

    What stands out:

    • It beats GPT-6 Luna on every benchmark where both have a score, at the same price. The gaps are largest on computer use (72.4% vs 48.9% on OSWorld) and agentic coding (39.2% vs 16.4% on Terminal-Bench 4.0).
    • The jump from Haiku 4.5 is huge. OSWorld goes from 15.7% to 72.4%, and Terminal-Bench goes from 0% to 39.2%. That's not a refresh, it's a different class of model.
    • Sonnet 5.5 still leads on hard agentic coding. 70.6% vs 39.2% on Terminal-Bench is a real gap. Anthropic says so directly: use Sonnet 5.5 or Opus 5.5 for complex agentic coding, and Haiku 5.5 for narrower tasks.
    • Effort matters a lot. According to VentureBeat, Haiku's Terminal-Bench score is about 39% at maximum effort and about 20% at the default medium. Benchmark headlines usually reflect the top setting, so check which level your evals use.

    These are vendor-reported numbers. Run your own evals before switching a production workload.

    Effort Levels: The Biggest Change for Developers

    Haiku 5.5 is the first Haiku model with the effort parameter. Effort controls how many tokens the model spends, including thinking, text and tool calls, so it's now your main control for quality, latency and cost on a single model.

    EffortWhen to use it on Haiku 5.5
    lowChat, short tool calls, simple high-volume requests (classification, routing, extraction). Cheapest and fastest.
    mediumThe default. Most work, including agentic coding.
    highKnowledge work, longer agent tasks, strict instruction following.
    xhigh / maxOnly where your evals show a gain. At this point, compare against Sonnet 5.5 on cost and speed.

    One thing to watch: Anthropic's docs say that at low, in long agent prompts, the model is more likely to skip a search, stop early, or skip a check. Use low for short, well-defined jobs, not long autonomous loops.

    Calling Claude Haiku 5.5 with the effort parameter (TypeScript)

    ts
    import Anthropic from "@anthropic-ai/sdk";
    
    const client = new Anthropic();
    
    const response = await client.messages.create({
      model: "claude-haiku-5-5",
      max_tokens: 1024,
      output_config: { effort: "low" },
      messages: [
        {
          role: "user",
          content: "Classify this support ticket as billing, bug, feature_request or other: 'I was charged twice this month.'",
        },
      ],
    });
    
    const text = response.content.find((b) => b.type === "text");
    console.log(text?.text);

    Python

    python
    import anthropic
    
    client = anthropic.Anthropic()
    
    response = client.messages.create(
        model="claude-haiku-5-5",
        max_tokens=1024,
        output_config={"effort": "low"},
        messages=[
            {"role": "user", "content": "Summarize this changelog in three bullet points: ..."}
        ],
    )
    
    for block in response.content:
        if block.type == "text":
            print(block.text)

    Turning thinking off

    Thinking is on by default and counts toward max_tokens, so leave room for it. For the fastest possible responses you can send thinking: { type: "disabled" }, but only at high effort or below. At xhigh or max it returns a 400 error. Also, with thinking disabled you can't change effort mid-conversation.

    ts
    const fast = await client.messages.create({
      model: "claude-haiku-5-5",
      max_tokens: 256,
      thinking: { type: "disabled" },
      output_config: { effort: "low" },
      messages: [{ role: "user", content: "Extract the order ID from: 'Order #A-99812 never arrived.'" }],
    });

    Using Claude Haiku 5.5 as a Subagent

    This is where Haiku 5.5 is most useful. The pattern: a bigger model (Opus 5.5 or Sonnet 5.5) plans and makes decisions, and it hands the repetitive, parallel work to many cheap, fast Haiku 5.5 calls.

    Anthropic's launch customers describe exactly this:

    • Rogo uses a Haiku 5.5 subagent to pull segment revenue from 10-K filings while a larger model builds the deck.
    • Cognition uses Haiku 5.5 as the sidekick in Devin Fusion. With Opus 5.5 leading, it reports a FrontierCode score of 66.2 at lower cost and latency.

    Here's a minimal orchestrator / subagent setup in TypeScript. Opus 5.5 splits the job, Haiku 5.5 handles each piece in parallel, and Opus 5.5 combines the results:

    ts
    import Anthropic from "@anthropic-ai/sdk";
    
    const client = new Anthropic();
    
    const textOf = (msg: Anthropic.Message) =>
      msg.content.flatMap((b) => (b.type === "text" ? [b.text] : [])).join("");
    
    // 1. Haiku 5.5 subagent: one narrow, repeatable job per call.
    async function extractFacts(document: string): Promise<string> {
      const msg = await client.messages.create({
        model: "claude-haiku-5-5",
        max_tokens: 2048,
        output_config: { effort: "low" },
        system: "Extract revenue, growth and risk factors as terse bullet points. No commentary.",
        messages: [{ role: "user", content: document }],
      });
      return textOf(msg);
    }
    
    // 2. Opus 5.5 orchestrator: reasons over the subagents' output.
    export async function analyze(documents: string[], question: string) {
      const facts = await Promise.all(documents.map(extractFacts));
    
      const msg = await client.messages.create({
        model: "claude-opus-5-5",
        max_tokens: 8192,
        messages: [
          {
            role: "user",
            content: `Question: ${question}\n\nExtracted facts:\n${facts.join("\n---\n")}`,
          },
        ],
      });
      return textOf(msg);
    }

    Why this works: each extraction prompt stays well under the 100K tier, every call shares a cacheable system prompt, and the expensive model only sees condensed facts instead of every raw document. In production, add a concurrency limit and retries around Promise.all. I covered those patterns in building agentic AI systems and agentic error recovery and observability.

    Haiku 5.5 vs Sonnet 5.5 vs Opus 5.5: Which Should You Use?

    Claude Haiku 5.5Claude Sonnet 5.5Claude Opus 5.5
    Best forHigh-volume, latency-sensitive work: classification, extraction, routing, summaries, compaction, live support, browser useThe best balance of speed and intelligenceLong-running agentic coding and knowledge work
    Price (in / out per MTok)From $0.10 / $0.50$2 / $10$4 / $20
    Default effortmediumhighmedium
    SpeedFastestFastModerate
    Context1M1M1M

    A simple rule: if the task is narrow and you run it many times, start with Haiku 5.5. If it needs long, autonomous, multi-step reasoning, use Sonnet 5.5 or Opus 5.5, and let Haiku be its subagent.

    Anthropic notes that Haiku 5.5 is its fastest model at standard speed. Opus models in Fast Mode are still faster.

    What Launch Customers Reported

    These come from Anthropic's announcement:

    • Asana: over 30% lower latency on task completions and up to 2.5x faster inference per agent turn.
    • HubSpot: 92.8% average on its CRM evaluation suite over three runs, the best it has seen from a small model.
    • AlphaSense: 0.84 vs 0.76 for Haiku 4.5 across 400 queries, a statistically significant gain.
    • Box: 11 points higher than Haiku 4.5 at about half the latency.

    Migrating from Claude Haiku 4.5 to 5.5

    The model ID swap is one line, but a few behaviors changed. Here's the checklist I'd run:

    1. 01Change the model ID from claude-haiku-4-5 to claude-haiku-5-5 (Bedrock: anthropic.claude-haiku-5-5).
    2. 02Set effort explicitly. The default is medium. For classification and extraction pipelines, test low first because it's cheaper and faster.
    3. 03Raise max_tokens. Thinking is on by default and counts toward it. A limit tuned for Haiku 4.5 can cut off answers.
    4. 04Re-measure token counts. The updated tokenizer uses more tokens for the same text, so update cost dashboards and any truncation logic that counts tokens.
    5. 05Keep prompts under 100K wherever you can. Crossing it multiplies the per-token price by 5x.
    6. 06Cache shared prefixes. Cache reads at $0.01 per million tokens make long, reused system prompts almost free.
    7. 07Check cybersecurity prompts. Haiku 5.5's cyber safeguards are stricter than Haiku 4.5's (but looser than Sonnet 5.5's). It allows more defensive work but still blocks penetration testing and attacker-style techniques. Teams with legitimate needs can apply to Anthropic's Cyber Verification Program or Life Sciences program.
    8. 08Run your evals at more than one effort level before you cut over. Anthropic's own numbers show big differences between medium and max.

    The Other Launch-Day Changes

    Sonnet 5.5 cache reads cut in half

    Sonnet 5.5 cache reads dropped from $0.20 to $0.10 per million tokens. Agent loops resend the same long context every turn, so Anthropic estimates this makes Sonnet 5.5 about 20% cheaper on most agentic workloads. You don't need to change any code to get it. According to VentureBeat, some Google Cloud and Azure customers get the cut a few days later.

    Monthly API credits for Max and Team

    Claude subscriptions now include API credits that work on any model through Anthropic's platform:

    PlanMonthly API credit
    Max 5x$100
    Max 20x$200
    TeamUp to $500, pooled

    If you pay for Max and also prototype against the API, that covers a lot of Haiku 5.5 calls: at $0.10 per million input tokens, $100 buys about a billion input tokens.

    Computer use and browser use in the SDKs

    The Python and TypeScript SDKs added beta support for computer use and browser use. Together with Haiku 5.5's 72.4% on OSWorld 2.1, that makes cheap browser automation agents practical.

    My Take

    Small models have usually forced a trade: cheap enough to run on everything, but not smart enough to trust. Haiku 5.5 removes most of that trade for narrow tasks. When a small model scores 72% on a computer-use benchmark and costs $0.10 per million input tokens, a lot of pipelines that used a mid-size model "just to be safe" can move down a tier.

    In AI commerce work like mine at Modelia, that's the high-volume layer: tagging product attributes, routing requests, summarizing catalog data, and checking structured output. That work runs thousands of times a day, and each call is simple. It's a good fit for Haiku 5.5 at low effort, with Sonnet 5.5 or Opus 5.5 kept for the steps that need real reasoning.

    The part I'd be careful with is effort. The benchmark headlines come from high effort settings, while the default is medium, and low can stop early on long agent tasks. Pick the effort level per task, measure it, and don't assume the launch numbers will carry over to your settings.

    If you're building on the Claude API, my guides on Claude agentic workflows in production and the Model Context Protocol pair well with this one. For more on the rest of Anthropic's 2026 lineup, see the Claude Opus 4.8 breakdown and the Claude Fable 5 relaunch.


    Sources: Anthropic: Claude Haiku 5.5, Claude models overview, Claude effort docs, VentureBeat launch coverage, 9to5Mac, and Daniela Amodei's LinkedIn post. Header image: public domain (CC0) via Wikimedia Commons.

    Harsh Rastogi is an AI Product Engineer at Modelia, building production generative-AI systems for fashion commerce, and the creator of carcode. He writes about AI systems, developer tooling and production engineering at harshrastogi.tech.

    Frequently asked questions

    What is Claude Haiku 5.5?

    Claude Haiku 5.5 is Anthropic's smallest and fastest model in the Claude 5.5 family, released on October 7, 2026. Anthropic calls it the cheapest, fastest and most capable small model it has released. It is built for high-volume, latency-sensitive work such as classification, extraction, routing, summaries and subagent tasks.

    How much does Claude Haiku 5.5 cost?

    For prompts up to 100K tokens, Claude Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens, with cache reads at $0.01 and cache writes at $0.125. Above 100K tokens it costs $0.50 input and $2.50 output. The Batch API is 50% off.

    What is the Claude Haiku 5.5 model ID?

    The Claude API model ID is claude-haiku-5-5. On Amazon Bedrock it is anthropic.claude-haiku-5-5, and Google Cloud and Microsoft Foundry use claude-haiku-5-5.

    What is the context window of Claude Haiku 5.5?

    Claude Haiku 5.5 has a 1M-token context window and up to 128K output tokens on the Messages API. Its reliable knowledge cutoff is June 2026.

    Is Claude Haiku 5.5 better than GPT-6 Luna?

    On Anthropic's published benchmarks, Haiku 5.5 scores higher than GPT-6 Luna on every benchmark where both have a result, including OSWorld 2.1 (72.4% vs 48.9%), Terminal-Bench 4.0 (39.2% vs 16.4%) and GDPval-AA (1620 vs 1437), at the same $0.10/$0.50 list price. These are vendor-reported numbers, so test on your own workload.

    Does Claude Haiku 5.5 support the effort parameter?

    Yes. Haiku 5.5 is the first Haiku model with adjustable effort. It supports low, medium, high, xhigh and max, and the default is medium. Set it with output_config.effort on the request.

    Should I use Haiku 5.5 or Sonnet 5.5?

    Use Haiku 5.5 for narrow, repeated tasks where speed and cost matter, such as classification, extraction, routing and subagents. Use Sonnet 5.5 or Opus 5.5 for complex, long-running agentic coding, where Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0 against Haiku 5.5's 39.2%.

    How is Claude Haiku 5.5 75% cheaper if the price per token is 90% lower?

    Per token, Haiku 5.5 is 90% cheaper than Haiku 4.5 for prompts up to 100K tokens. Its updated tokenizer produces somewhat more tokens for the same text, so Anthropic estimates the average real workload costs about 75% less.

    • Claude
    • Anthropic
    • Claude Haiku 5.5
    • LLM Pricing
    • AI Agents
    • Developer Tools

    Written by Harsh Rastogi, AI Product Engineer and Business AI Head at Modelia. More on AI products, agents and production engineering on LinkedIn.

    Share

    LinkedInX
    Portrait of Harsh Rastogi, AI Product Engineer

    Harsh Rastogi

    AI Product Engineer and Business AI Head at Modelia

    I build AI products and agents from the first sketch to production at Modelia, and I'm a Founding Engineer at SelfAgentic. Before that I built platforms at Asynq and Bharat Electronics Limited. I publish the agent skills I use every day, and you can see what I've shipped in my work.

    • Claude Fable 5 Is Back: Export Controls Lifted and Global Access Restored July 1

      On June 30, 2026 the US lifted the export controls it placed on Claude Fable 5 and Mythos 5, and Fable 5 came back online globally on July 1. Here's the full timeline, what changed under the hood — a 99%+ bypass classifier, refusals, and fallback billing — and what developers need to do to ship on claude-fable-5.

      • AI
      • Anthropic
      • Claude
      • Developer Tools
      AI & Machine Learning9 min read
    • Model Context Protocol (MCP): A Production Engineer's Complete Guide

      Model Context Protocol hit 97 million installs in 2026 but most tutorials stop at hello-world. A hands-on guide to building production MCP servers with real code, auth, tool registration, and observability — from a production engineer at Modelia.

      • MCP
      • Anthropic
      • Claude
      • Agentic AI
      AI & Machine Learning12 min read

    Get the next article by email.

    New articles and AI Pulse editions. Nothing else.