Claude Haiku 5.5 Pricing: The 100K Line Behind the 90% Cut

· AIHubMix · 8 min read · Opinion

Claude Haiku 5.5 Pricing: The 100K Line Behind the 90% Cut

Claude Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens, a tenth of Haiku 4.5's $1 and $5. That holds for prompts up to 100,000 tokens. Above that line, the whole request is billed at $0.50 and $2.50, which is half of Haiku 4.5's price rather than a tenth.

Anthropic puts the typical saving at around 75%, not 90%, and the difference comes from three things this post works through: where your prompts fall relative to the 100K line, how many thinking tokens the default settings generate, and a new tokenizer that counts the same text as about 30% more tokens. Two worked examples below show how a real bill comes out.

The rate card

Prices per million tokens, from Anthropic's Haiku 5.5 launch post and pricing page:

Token type Haiku 5.5, short Haiku 5.5, long Haiku 4.5 Sonnet 5.5
Input $0.10 $0.50 $1.00 $2.00
Output $0.50 $2.50 $5.00 $10.00
Cache read $0.01 $0.05 $0.10 $0.10
Cache write, 5 min $0.125 $0.625 $1.25 $2.50
Cache write, 1 hour $0.20 $1.00 $2.00 $4.00

"Short" means a prompt of 100,000 tokens or fewer; "long" means over 100,000. The Batch API takes 50% off input and output. Sonnet 5.5's cache read price was cut from $0.20 to $0.10 on the same day.

The Haiku 5.5 page on AIHubMix lists the same two tiers, with web search at $0.01 per request.

Billing rules that are easy to miss

The whole request changes tier, not just the excess. Anthropic prices Haiku 5.5 with two rate cards chosen by prompt length. A prompt of 100,000 tokens pays $0.0100 for its input; a prompt of 100,001 tokens pays $0.0500, and its output costs five times as much too. Haiku 5.5 is the exception here: Anthropic's other current 1M-context models bill the full window at their standard rates.

The line arrives sooner than your old dashboards suggest. Haiku 5.5 uses a newer tokenizer, and the same text produces about 30% more tokens than on Haiku 4.5. Dividing 100,000 by 1.3 puts the line near 77,000 tokens as Haiku 4.5 counted them. A 78,000-token prompt on your old dashboard becomes roughly 101,000 tokens on Haiku 5.5, and pays the long rate. That arithmetic comes from the 30% figure in Anthropic's migration guide; the exact increase depends on the content, so recount real prompts on the new model.

Thinking is on by default, and it bills as output. Haiku 4.5 only thought when asked. Haiku 5.5 runs adaptive thinking at medium effort unless you change it, so a route that used to return 200 output tokens can now return several hundred more.

Short prompts can now be cached. The minimum cacheable prompt drops from 4,096 tokens on Haiku 4.5 to 512 on Haiku 5.5, according to Anthropic's migration guide. A 3,000-token system prompt that could never be cached on Haiku 4.5 can be cached now, and a cache read costs a tenth of the input rate.

Priority Tier is not available. If you hold a Priority Tier commitment on Haiku 4.5, it does not carry over to Haiku 5.5. Plan capacity separately.

Cached tokens and the 100K line: assume they count. Anthropic lists a separate cache read price for each tier, which suggests the whole prompt, cached or not, decides the tier. Anthropic has not spelled this out, so the safe assumption is that a 95K cached prefix plus a 6K question is a long prompt.

Worked example 1: a support-ticket classifier

Assumptions, in Haiku 4.5 token counts: a 3,000-token system prompt with tool definitions, a 1,000-token ticket, a 200-token answer, and 1 million requests a month. Haiku 4.5 runs with thinking off.

On Haiku 5.5 the same text becomes 3,900 + 1,300 input tokens and a 260-token answer. Thinking is assumed to add 300 tokens at low effort and 800 at medium. Those thinking figures are assumptions; measure yours.

Setup Per request Per month
Haiku 4.5 $0.00500 $5,000
Haiku 5.5, low, no cache $0.00080 $800
Haiku 5.5, low, cached prefix $0.00045 $453
Haiku 5.5, medium, cached prefix $0.00070 $703
Sonnet 5.5, medium, cached prefix $0.01359 $13,674

Monthly cached rows include about $4 of cache writes for Haiku 5.5 and about $84 for Sonnet 5.5, assuming one 5-minute write every five minutes. Sonnet 5.5 uses the same token assumptions as Haiku at medium.

How the cached low row adds up: 3,900 cached tokens at $0.01 per million ($0.000039), plus 1,300 fresh input tokens at $0.10 ($0.00013), plus 560 output tokens at $0.50 ($0.00028), for $0.000449 per request.

The pattern: caching and effort together move the bill from 16% of the old cost (no cache, low) to 9% (cached, low). Leaving effort at the default instead of low adds about $250 a month at this volume. Neither choice matters much in absolute terms next to Haiku 4.5's $5,000, which is why the classifier case is where the 90% headline is real.

Worked example 2: a contract review that crosses the line

Assumptions, in Haiku 4.5 token counts: a contract plus instructions of 70,000 tokens, a 1,500-token answer, no caching. On Haiku 5.5 that becomes 91,000 input tokens, a 1,950-token answer, and an assumed 2,000 thinking tokens at medium, so 3,950 output tokens.

Now make the contract 14% longer: 80,000 tokens on Haiku 4.5, 104,000 on Haiku 5.5.

Contract size (Haiku 4.5 count) Haiku 4.5 Haiku 5.5 Change
70,000 tokens $0.0775 $0.0111 86% lower
80,000 tokens $0.0875 $0.0619 29% lower

The second row: 104,000 input tokens at $0.50 per million ($0.052) plus 3,950 output tokens at $2.50 ($0.0099). Fourteen percent more text costs 5.6 times as much on Haiku 5.5.

The fix is to stay under the line. Splitting the longer contract into two calls of about 53,000 tokens each (half the contract plus the instructions in each) costs about $0.015 in total at the short rate, under a quarter of the single long call. Whether a split works depends on the task; clause-by-clause review splits cleanly, a question that needs the whole document at once doesn't.

How it compares on cost per task

Against Sonnet 5.5, Haiku 5.5 is about a twentieth of the per-token price on short prompts. But price per token is not price per finished task. On complex agentic coding, Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0 to Haiku 5.5's 39.2%, in Anthropic's reported results, and a cheaper request that fails and gets retried, or gets redone by a larger model, is not cheaper. Anthropic's own advice is to compare against Sonnet 5.5 before running Haiku at xhigh or max. Sonnet 5.5 also got cheaper the same day: its cache reads were halved, which Anthropic says cuts about 20% off most agentic work.

Against GPT-6 Luna, the short-prompt list prices are identical, cache reads included. The difference is the long-prompt rule. Luna's long-context rate starts above 272,000 input tokens and is $0.20 / $0.75, while Haiku 5.5's starts above 100,000 and is $0.50 / $2.50. For a workload between 100K and 272K tokens, Luna is far cheaper on list price. The two models use different tokenizers, so compare measured bills, not token counts.

Figures on the 100K threshold and the tokenizer effect have also been worked through by Roo's Newsletter, whose arithmetic matches the examples above.

Steps that lower the bill

  1. Count tokens on the new model and alert before 100K. Log the full prompt size for every request and alert at around 90,000 tokens, so prompts that drift upward get trimmed or split before they change tier.
  2. Cache the stable prefix. System prompt and tool definitions first, variable content last. With a 512-token minimum, most production system prompts now qualify. AIHubMix's Claude prompt caching guide shows the request format.
  3. Set effort explicitly. Use low for high-volume, simple routes. The default medium costs more on every request whether the route needs it or not.
  4. Size max_tokens for thinking plus answer. Too small and you pay for thinking that ends with no answer.
  5. Batch what can wait. The Batch API halves input and output rates.
  6. Route up instead of retrying down. If a task type fails often on Haiku, send it to Sonnet 5.5 first rather than paying for Haiku attempts plus a fallback.

A logging pattern for steps 1 and 2 through AIHubMix's Claude native endpoint:

import os
import anthropic

client = anthropic.Anthropic(
    api_key=os.environ["AIHUBMIX_API_KEY"],
    base_url="https://aihubmix.com",
)

SYSTEM_PROMPT = "You triage support tickets..."  # stable, cacheable prefix

def classify(ticket: str) -> str:
    r = client.messages.create(
        model="claude-haiku-5-5",
        max_tokens=1024,
        output_config={"effort": "low"},
        system=[{"type": "text", "text": SYSTEM_PROMPT,
                 "cache_control": {"type": "ephemeral"}}],
        messages=[{"role": "user", "content": ticket}],
    )
    u = r.usage
    prompt_tokens = (u.input_tokens + (u.cache_read_input_tokens or 0)
                     + (u.cache_creation_input_tokens or 0))
    if prompt_tokens > 90_000:
        print(f"warning: prompt at {prompt_tokens} tokens, near the 100K price line")
    return next(b.text for b in r.content if b.type == "text")

If cache_read_input_tokens stays at zero across repeated calls, something in the prefix is changing between requests, such as a timestamp in the system prompt.

FAQ

How much does Claude Haiku 5.5 cost?
For prompts up to 100,000 tokens, $0.10 per million input tokens and $0.50 per million output tokens. For longer prompts, $0.50 and $2.50, applied to the whole request. Cache reads cost a tenth of the input rate, and the Batch API halves input and output prices.

Is Haiku 5.5 really 90% cheaper than Haiku 4.5?
Per token, on short prompts, yes. Per task, less: the new tokenizer counts about 30% more tokens for the same text, thinking is on by default, and long prompts get only a 50% cut. Anthropic estimates a typical saving of about 75%. A short-prompt classifier with caching can save around 90%.

What happens if my prompt is just over 100,000 tokens?
The entire request moves to the higher rate card, input and output alike. In the contract example above, 14% more text made the request 5.6 times as expensive.

Do cached tokens count toward the 100K threshold?
Anthropic hasn't stated this explicitly. Its pricing lists separate cache read prices for each tier, so the safe assumption is that the whole prompt, cached or not, decides the tier.

Do thinking tokens cost extra?
They are billed as output tokens at the normal output rate. Because thinking is on by default at medium effort, they can add noticeably to a route that used to produce short answers. Lowering effort reduces them.

Is Haiku 5.5 cheaper than GPT-6 Luna?
On prompts up to 100K tokens the list prices are the same. Between 100K and 272K tokens, Luna stays at its base rate while Haiku 5.5 moves to its higher one, so Luna is cheaper there on list price. Tokenizers differ, so compare real bills.

Can I use Priority Tier with Haiku 5.5?
No. Priority Tier isn't supported on Haiku 5.5, so commitments on Haiku 4.5 don't transfer.

Keep reading: the Claude Haiku 5.5 series

Sources

More from the blog