Claude Haiku 5.5 Effort Levels: Medium Is the Default, Low Is Often Enough

· AIHubMix · 8 min read · Tutorial

Claude Haiku 5.5 Effort Levels: Medium Is the Default, Low Is Often Enough

Claude Haiku 5.5 is the first Haiku with an effort setting, and the setting moves the bill more than anything else you control. In Artificial Analysis's independent runs, Haiku 5.5 at max effort scored 43 on its Intelligence Index and at low effort scored 29. Max also used 440 million output tokens to get through the index; low used 32 million.

So the short answer: leave most work at the default, medium. Drop high-volume, simple routes to low. Raise knowledge work and strict instruction following to high. Treat xhigh and max as settings you need evidence for, and compare them against Sonnet 5.5 before you commit.

The rest of this post shows what each level costs, where moving up stops paying off, and which settings break or get expensive when you change effort.

What the setting is

Claude Haiku 5.5 supports five levels: low, medium, high, xhigh, and max. The default is medium, which is unusual: most current Claude models default to high, and only Haiku 5.5 and Opus 5.5 default one level lower. Sending medium is the same as leaving the parameter out.

Effort replaces the manual thinking budget that Haiku 4.5 used. A request with thinking: {"type": "enabled", "budget_tokens": N} now returns a 400, so there is no old number to carry over. Per Anthropic's effort documentation, you set it in output_config:

output_config={"effort": "low"}

Three facts frame everything below:

  • Thinking is on by default. Haiku 5.5 runs adaptive thinking unless told otherwise. Lower effort means less thinking, and on simple requests the model can skip thinking entirely.
  • Thinking tokens are output tokens. They count toward max_tokens and bill at the output rate ($0.50 per million on prompts up to 100K tokens).
  • Asking for less thinking in the prompt doesn't work. In Anthropic's testing, the model still thought when the prompt told it to answer directly. Effort is the lever.

What each level costs

Two independent sources measured Haiku 5.5 across effort levels. Index figures come from Artificial Analysis; the FrontierCode figures come from Cognition's public result file, as compiled by Kingy.ai. Neither covers your workload, so treat them as the shape of the curve, not as your numbers.

Artificial Analysis Intelligence Index v4.3.2 (ten evaluations, from knowledge work to terminal tasks):

Effort Index score Output tokens for the index Cost per task Time to first token
low 29 32M $0.02 9.9 s
medium 34 54M $0.05 13.4 s
high 38 97M $0.08 28.0 s
xhigh 41 180M $0.12 n/a
max 43 440M $0.21 n/a

Artificial Analysis has not published a latency figure for xhigh. Its max-effort latency figure looked anomalous, so it is left out.

FrontierCode 1.1 Main, Haiku 5.5 running in Claude Code:

Effort Composite score Cost per rollout Output tokens per rollout
low 34.8% $0.06 23,868
medium 41.6% $0.13 36,172
high 41.9% $0.26 55,149
xhigh 45.8% $0.63 100,847
max 46.4% $1.33 181,387

Costs are rounded to the cent.

Read the two tables together and three things stand out.

Low to medium is the cheapest step up. On the index, it buys 5 points for about 1.7 times the output tokens. On FrontierCode, it buys 6.8 points for about twice the cost per rollout.

Medium to high can be flat. On FrontierCode, high scored 0.3 points above medium and cost about twice as much. On the broader index, high did add 4 points. Whether your workload looks like the first case or the second is exactly what a quick eval tells you.

The top two levels are expensive. From high to max on the index, output tokens rise about 4.5 times for 5 points. On FrontierCode, max costs about ten times what medium costs for 4.8 more points, and the step from xhigh to max roughly doubles cost for 0.6 points. Anthropic's own guidance says the same thing in plainer terms: use xhigh and max only where your evals show a gain, and compare them with Sonnet 5.5 on performance, cost, and speed first.

One more reason the default matters: Anthropic's launch benchmarks were run at max effort. The system card's medium-effort results are lower (1277 on GDPval-AA instead of 1620). If you deploy at the default, those are the numbers to compare against.

Where to start, by workload

Workload Start at Move up when
Classification, routing, tagging low Labels drift on hard cases
Extraction from short documents low Fields are missed or merged
Chat and live support low Rules slip in long chats
Summaries and compaction medium Key details drop out
Subagent work under a larger model medium The lead model redoes the work
Agentic coding, narrow changes medium Tests fail on first attempt
Knowledge work, long agent tasks high Evals show headroom

These starting points follow Anthropic's guidance for Haiku 5.5; the "move up" signals are what to watch for in your own logs.

The one row worth a second look is chat. Low is the fastest level, which suits live support. But Anthropic recommends high when instruction following matters most, for example a support assistant that must hold its system-prompt rules while a user argues. Test both on the conversations that went wrong in the past.

What breaks or gets expensive when effort changes

A small max_tokens can end the reply before it starts. Thinking counts toward max_tokens. A cap sized for a one-word classification, say 50 tokens, can be spent entirely on thinking, and the response ends with stop_reason: "max_tokens" and no text. Raise the cap to leave room, or lower effort.

Changing effort mid-conversation resets the cache. Changing the top-level effort value between requests invalidates the prompt cache for the conversation's messages. If an agent needs one hard turn at high inside a low conversation, Anthropic offers a per-message effort change (beta header mid-conversation-output-config-2026-07-01, Claude API and Google Cloud) that keeps the cache. It requires thinking to be on. Whether a gateway passes that beta header through is worth checking before you rely on it.

Turning thinking off has limits. thinking: {"type": "disabled"} is accepted at low, medium, and high. At xhigh or max it returns a 400. With thinking off, a per-message effort change also returns a 400, and the model may skip a tool call it needs when the same request asks for JSON output. Anthropic's advice is to keep thinking on and use a lower effort level instead.

Low effort can stop early. With a long coding-agent system prompt at low, Haiku 5.5 sometimes stops before the work is done and hands the task back. In Anthropic's testing, moving from low to medium roughly halved early stopping and more than doubled output tokens per attempt. The Haiku 5.5 prompting guide has a short "keep working until done" instruction that addresses the same problem at the lower price.

Low and medium can skip verification. On coding tasks, the model sometimes reports a change as done without running a test. The prompting guide has a verification paragraph for this; it costs tokens but raises the share of checked changes.

Xhigh can return an empty reply. In multi-turn chats at xhigh, the model sometimes writes its whole answer inside its thinking and ends the turn with no visible text. If you run at xhigh, check for empty text blocks.

Forcing a tool skips thinking. Haiku 5.5 accepts a forced tool_choice, but the response then starts with the tool call and no thinking. If the model needs to reason before calling the tool, use auto and say in the prompt when to call it.

Run your own effort sweep through AIHubMix

The fastest way to choose is to run the same 20 to 50 real prompts at two or three levels and compare quality, output tokens, and latency. Through the AIHubMix Claude native endpoint, with AIHUBMIX_API_KEY set:

import os
import time
import anthropic

client = anthropic.Anthropic(
    api_key=os.environ["AIHUBMIX_API_KEY"],
    base_url="https://aihubmix.com",
)

prompts = [
    "Summarize this support ticket in one sentence: ...",
    "Extract the invoice number and total from: ...",
]

for effort in ["low", "medium", "high"]:
    out_tokens, seconds = 0, 0.0
    for prompt in prompts:
        start = time.time()
        r = client.messages.create(
            model="claude-haiku-5-5",
            max_tokens=8000,
            output_config={"effort": effort},
            messages=[{"role": "user", "content": prompt}],
        )
        seconds += time.time() - start
        out_tokens += r.usage.output_tokens
        text = next((b.text for b in r.content if b.type == "text"), "")
        # Save `text` next to the prompt and grade it against your own rubric.
    print(f"{effort}: {out_tokens} output tokens, {seconds:.1f}s total")

If output tokens barely change between low and high, the setting is probably not reaching the model; check the request your gateway forwards. On the Haiku 5.5 page on AIHubMix, the per-token price is the same at every effort level, so the token count from this sweep converts straight into dollars.

FAQ

What is the default effort level for Claude Haiku 5.5?
Medium. Most current Claude models default to high, so code that relied on the default for another model will run Haiku 5.5 one level lower unless it sets effort explicitly.

Can I turn thinking off completely?
At low, medium, and high effort, yes, by setting thinking to disabled. At xhigh and max that returns an error. Anthropic recommends lowering effort instead, because the model can skip thinking on simple requests by itself and keeps its tool-calling behavior intact.

Which effort level is cheapest?
Low. In Artificial Analysis's runs it used about 32 million output tokens for the whole index against 54 million at medium and 440 million at max. The per-token price is the same at every level; the difference is how many tokens get generated.

Is max effort worth it on Haiku 5.5?
Only with evidence from your own evals. On FrontierCode, max cost about ten times as much as medium for a 4.8-point gain. At that point Anthropic suggests comparing against Sonnet 5.5, which may reach the same quality for less.

Why do my responses stop with no text?
Usually because thinking used up max_tokens before the answer began. Raise max_tokens or lower effort. At xhigh in multi-turn chats, an empty reply can also mean the model put its answer inside its thinking.

Does changing effort affect prompt caching?
Yes. Changing the top-level effort between requests invalidates the cached messages for that conversation. A per-message effort change, currently in beta, avoids that.

Why don't the launch benchmarks match what I see at the default?
The launch figures were run at max effort. At medium, the system card reports lower scores, for example 1277 instead of 1620 on GDPval-AA.

Keep reading: the Claude Haiku 5.5 series

Sources

More from the blog