Claude Haiku 5.5 vs Claude Sonnet 5.5
Sonnet costs 20x more per token — when is it actually worth it?
The short answer
| At a glance | Claude Haiku 5.5 | Claude Sonnet 5.5 |
|---|---|---|
| List priceinput / output per 1M tokens | $0.10 / $0.50 (better) | $2 / $10 |
This is a volume-vs-difficulty decision, not a coin flip. Per token, Sonnet 5.5 is 20x Haiku's ≤100K price ($2/$10 vs $0.10/$0.50) — and above 100K input tokens the gap narrows to 4x because Haiku jumps to its $0.50/$2.50 tier while Sonnet stays flat. On quality, the only same-run public data (Anthropic's Haiku-launch table) has Sonnet clearly ahead on every benchmark, but the gap is task-shaped: close-ish on computer use (OSWorld 83.9% vs 72.4%) and huge on agentic coding (Terminal-Bench 4.0: 70.6% vs 39.2%). The working rule: route classification, extraction, routing, and high-volume chat to Haiku (Anthropic's own positioning for it), send hard reasoning, repo-level coding, and long-horizon agents to Sonnet — and watch Haiku's tokenizer caveat ('slightly more tokens per task'), which quietly narrows the price gap on real workloads. One more wrinkle: Sonnet 5.5 lists at exactly GPT-6.1 Sol's price, so cross-vendor options belong in the decision too.
Choose Claude Haiku 5.5
High-volume, latency-sensitive work: classification, extraction, routing, support chat — the use cases Anthropic explicitly built it for. 20x cheaper per token below 100K, and 'our fastest model to date' at standard speed.
Choose Claude Sonnet 5.5
Hard tasks where the same-run quality gap is large: agentic coding (Terminal-Bench 70.6% vs 39.2%), long-horizon agents, prompts over 100K tokens (Haiku's 5x tier erodes its price advantage), or when output quality errors cost more than tokens.
Either — decide by evidence you generate
Mixed fleets: the honest answer is routing. Most real products use Haiku for the 90% easy traffic and Sonnet (or Opus) for the 10% hard traffic — model the split in the calculator before committing.
Key facts
- List prices: Haiku 5.5 $0.10/$0.50 per MTok up to 100K input tokens per request, $0.50/$2.50 above; Sonnet 5.5 $2/$10 flat at any size (both verified 2026-10-09).
- The 20x/4x structure: below 100K Sonnet is 20x Haiku on both rates; above 100K Haiku's tier makes it 4x — big-prompt workloads hurt Haiku's economics twice (5x rates AND slightly more tokens per task from the new tokenizer).
- Same-run quality (Anthropic's Haiku-launch table, both models in one run): Terminal-Bench 4.0 70.6% vs 39.2%; OSWorld 2.1 83.9% vs 72.4%; GDPval-AA 1840 vs 1620; AA-Briefcase 1824 vs 1578; FrontierCode 52.1% vs 46.4% (Sonnet xhigh); Chartography 61.6% vs 46.4%.
- Haiku 5.5 massively closed the gap on its predecessor: Terminal-Bench 4.0 went 0.0% (Haiku 4.5) to 39.2%. 'Good enough' moved a lot — but not to Sonnet's level on hard agentic tasks.
- Cache economics: Sonnet reads at 5% of base ($0.10); Haiku reads at 10% of its base ($0.01) — both effectively $0.01–$0.10, so cache-heavy loops don't change the ranking.
- Neither model has a fast mode (Opus-only perk per Anthropic's pricing page); both get Batch at −50%; both support 1M context and adaptive thinking — but default effort differs: Sonnet ships at high, Haiku at medium.
- Cross-vendor check: Sonnet 5.5 costs exactly what GPT-6.1 Sol costs ($2/$10) — if you're choosing a mid-tier model, that comparison belongs on your shortlist too.
- Retirement: Haiku not sooner than 2027-10-07, Sonnet not sooner than 2027-09-28 — both safe horizons for production planning.
Official pricing, every tier
USD per 1M tokens from vendor pricing pages. 'Not published' means exactly that — we don't extrapolate.
| Service tier | Claude Haiku 5.5 | Claude Sonnet 5.5 |
|---|---|---|
| Standarddefault on-demand serving | in $0.10 · cached $0.01 · out $0.50 cache write $0.125 · >100K in/req | in $2 · cached $0.10 · out $10 cache write $2.5 |
| Batchasync jobs, 50% off | in $0.05 · cached $0.005 · out $0.25 cache write $0.0625 · derived · >100K in/req | in $1 · cached $0.05 · out $5 cache write $1.25 · derived |
| Long-context band | Claude Haiku 5.5 | Claude Sonnet 5.5 |
|---|---|---|
| Long-context ratesapplies per request above the model threshold | in $0.50 · cached $0.05 · out $2.5 cache write $0.625 · >100K in/req | Flat pricing same rates at any input size (up to context window) |
USD per 1M tokens. Every input token is billed in exactly one bucket — uncached, cached read, or cache write; a write replaces the input rate for those tokens (official rule), so write rates are per tier as listed. Anthropic rows bill the 5-minute cache-write TTL; the 1-hour TTL ($8/MTok standard) is also published, and cache multipliers stack with the Batch discount and Fast-mode pricing. · Anthropic pricing (verified 2026-10-09)
Tier ≠ quality change
Run this comparison on your own numbers
Preloaded with a workload typical for this comparison. Change anything — results recompute locally.
| Model | Per request | Monthly |
|---|---|---|
| Claude Haiku 5.5 cheapest | $0.00027 | $540.00 |
| Claude Sonnet 5.5 | $0.0052 | $10,400.00 |
Monthly spread between the cheapest and priciest selected model: $9,860.00.
Breakdown: Claude Haiku 5.5 · standard tier · standard context
- Fresh input (uncached): 1,000 tok × $0.1/M$0.0001
- Cached input (reads): 2,000 tok × $0.01/M$0.00002
- Cache writes: 0 tok × $0.125/M$0.00
- Output (incl. reasoning tokens): 300 tok × $0.5/M$0.00015
- Per request$0.00027
Rates: official pricing · verified 2026-10-09
- Per-token arithmetic only: it does not model quality differences or cross-vendor output-token-count differences for the same task.
- Reasoning tokens are billed as output tokens at the chosen model's output rate.
- Long-context band is decided per request by total input tokens vs the provider threshold, not by monthly volume.
- Cache accounting uses the official mutually-exclusive buckets: each input token is billed once, as uncached input, cached read, OR cache write — writes replace the input rate for those tokens, never add to it.
- Claude rows bill the 5-minute-TTL cache write; the published 1-hour TTL ($8/MTok standard) is not modeled (no TTL input).
- Rows flagged 'derived' are computed from a stated official rule (e.g. Anthropic batch -50%), not read verbatim from a price table.
- Missing official prices render as Unsupported — we never apply a guessed multiplier (e.g., GPT-6 Sol has no Ultrafast tier).
What the evidence actually says
Independent and vendor-run evals, each with publisher, runner, and harness. Numbers from different runs are never merged.
| Benchmark | Score | Effort | Harness | Evidence |
|---|---|---|---|---|
| GDPval-AAv2.1⚠ run comparability unknown | ||||
| claude-haiku-5-5 | 1620 elo | unspecified | Anthropic Haiku 5.5 announcement evals (GDPval-AA benchmark by Artificial Analysis) | Vendor-reported |
| claude-sonnet-5-5 | 1840 elo | — | Anthropic Haiku 5.5 announcement evals (GDPval-AA benchmark by Artificial Analysis) | Vendor-reported |
| claude-haiku-5-5 | 1618 elo | — | Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models) | Third-party |
| AA-Briefcasev1.1⚠ run comparability unknown | ||||
| claude-haiku-5-5 | 1578 elo | unspecified | Anthropic Haiku 5.5 announcement evals | Vendor-reported |
| claude-sonnet-5-5 | 1824 elo | — | Anthropic Haiku 5.5 announcement evals | Vendor-reported |
| claude-haiku-5-5 | 1577 elo | — | Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models) | Third-party |
| OSWorld2.1 offline subset | ||||
| claude-haiku-5-5 | 72.4 % | unspecified | Anthropic Haiku 5.5 announcement evals | Vendor-reported |
| claude-sonnet-5-5 | 83.9 % | — | Anthropic Haiku 5.5 announcement evals | Vendor-reported |
| Terminal-Bench4.0⚠ run comparability unknown | ||||
| claude-haiku-5-5 | 39.2 % | unspecified | Anthropic Haiku 5.5 announcement evals | Vendor-reported |
| claude-haiku-5-5 | 35.35 % | max | Vals AI (independent) | Third-party |
| claude-sonnet-5-5 | 70.6 % | — | Anthropic Haiku 5.5 announcement evals | Vendor-reported |
| claude-haiku-5-5 | 33 % | — | Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models) | Third-party |
| FrontierCode1.1 (Main) | ||||
| claude-haiku-5-5 | 46.4 % | unspecified | Anthropic Haiku 5.5 announcement evals | Vendor-reported |
| claude-sonnet-5-5 | 52.1 % | xhigh | Anthropic Haiku 5.5 announcement evals | Vendor-reported |
| Chartographyunspecified (no tools) | ||||
| claude-haiku-5-5 | 46.4 % | unspecified | Anthropic Haiku 5.5 announcement evals | Vendor-reported |
| claude-sonnet-5-5 | 61.6 % | — | Anthropic Haiku 5.5 announcement evals | Vendor-reported |
| Vals Index (overall)unspecified | ||||
| claude-haiku-5-5 | 54.31 % | max | Vals AI (independent, no refusal fallback) | Third-party |
| AA Intelligence Indexv4.3.2 | ||||
| claude-haiku-5-5 | 43 index-points | — | Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models) | Third-party |
| AutomationBench-AAunspecified | ||||
| claude-haiku-5-5 | 35 % | — | Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models) | Third-party |
| AA cost per Intelligence-Index taskv4.3.2 | ||||
| claude-haiku-5-5 | 0.21 usd-per-task | — | Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models) | Third-party |
Same benchmark name, multiple rows — run comparability unknown
GDPval-AA: claude-haiku-5-5 = 1620 (Anthropic Haiku 5.5 announcement evals (GDPval-AA benchmark by Artificial Analysis)); claude-sonnet-5-5 = 1840 (Anthropic Haiku 5.5 announcement evals (GDPval-AA benchmark by Artificial Analysis)); claude-haiku-5-5 = 1618 (Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models)). Same table contains the GPT-6 Luna row (1437) — a same-table, same-run comparison of exactly the pair our haiku-vs-luna page covers. Column order in the source table is Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | Sonnet 5.5; 735 belongs to Haiku 4.5 and MUST NOT be read as Luna's value (this exact column shift was a published error, corrected 2026-10-10). Different table from the Opus 5.5 announcement's GDPval (Opus 1846 / Astra 1542): never merge across tables.
AA-Briefcase: claude-haiku-5-5 = 1578 (Anthropic Haiku 5.5 announcement evals); claude-sonnet-5-5 = 1824 (Anthropic Haiku 5.5 announcement evals); claude-haiku-5-5 = 1577 (Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models)). Whether these rows share a run is not documented in the sources; treat them as not comparable.
Terminal-Bench: claude-haiku-5-5 = 39.2 (Anthropic Haiku 5.5 announcement evals); claude-haiku-5-5 = 35.35 (Vals AI (independent)); claude-sonnet-5-5 = 70.6 (Anthropic Haiku 5.5 announcement evals); claude-haiku-5-5 = 33 (Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models)). Vals AI independently measured Haiku 5.5 at 35.35% on Terminal-Bench 4.0 (see terminalbench4-haiku55-vals) — vendor vs independent runs differ by ~4 pp; both stored, neither merged.
We never average or merge these. See methodology.Sources: Anthropic — Introducing Claude Haiku 5.5 (2026-10-07) · Vals AI (independent) — Vals AI evaluation of Claude Haiku 5.5 (posted 2026-10-08) · Artificial Analysis (independent) — Independent same-harness comparison: Claude Haiku 5.5 vs GPT-6 Luna (Intelligence Index v4.3.2)
Which model for which workload
Confidence labels: Measured = same-harness evidence · Inferred = reasoned from documented evidence · Insufficient evidence = no comparable public data.
| Workload | Our pick | Confidence | Why |
|---|---|---|---|
| Daily coding tasks | Claude Sonnet 5.5 | Measured | Same-run FrontierCode 52.1% vs 46.4% and Terminal-Bench 70.6% vs 39.2%: routine coding leans Sonnet, sharply on agentic tasks. Vendor-run table; Sonnet at xhigh for FrontierCode. For light code Q&A, Haiku's 20x price advantage may dominate — know your task mix. |
| Agent workflows | Claude Sonnet 5.5 | Measured | OSWorld 2.1: 83.9% vs 72.4% same-run; GDPval-AA 1840 vs 1620 on professional agent work. Vendor-run; for high-volume simple routing agents, Haiku's price can still win the fleet economics. |
| Long-context tasks | Claude Sonnet 5.5 | Price-derived | Sonnet is flat $2/$10 at any size; Haiku's >100K tier ($0.50/$2.50) plus tokenizer overhead makes long-prompt workloads relatively MORE expensive on the small model. |
| Budget-sensitive / high volume | Claude Haiku 5.5 | Price-derived | 20x cheaper below 100K per token — classification/extraction/routing volume work is what Anthropic built it for. Official tokenizer caveat ('slightly more tokens per task') means realized savings sit below 20x. |
| Quality-first | Claude Sonnet 5.5 | Measured | Ahead on every same-run benchmark; if quality still isn't enough, Opus 5.5 ($4/$20) is the next step up. Vendor-run table. |
| Latency-sensitive production | Decide by bake-off | Insufficient evidence | Conflicting vendor positioning: Anthropic calls Haiku 'our fastest model to date' while listing Sonnet's latency as 'Fast' with 'best combination of speed and intelligence'. No independent same-protocol comparison exists. |
| Offline batch jobs | Claude Haiku 5.5 | Price-derived | Batch at −50% keeps the 20x gap ($0.05/$0.25 vs $1/$5); quality-tolerant bulk work belongs on Haiku. |
| Complex repository maintenance | Claude Sonnet 5.5 | Measured | Terminal-Bench's 70.6%-vs-39.2% gap is exactly the repo-scale agentic-coding profile; also consider Opus 5.5 or GPT-6.1 Sol at the same price. |
Frequently asked questions
What does Claude Sonnet 5.5 cost?
$2 per million input tokens and $10 per million output tokens, flat at any prompt size — cache reads $0.10 (5% of base), 5-minute cache writes $2.50, Batch −50%. No fast mode, no prompt-size tiers. That's the same list price as OpenAI's GPT-6.1 Sol and GPT-6 Sol.
Is Sonnet 5.5 worth 20x the price of Haiku 5.5?
Depends entirely on task difficulty. On the same-run table Sonnet leads everywhere, but the gaps split: computer use 83.9% vs 72.4% (close), Terminal-Bench agentic coding 70.6% vs 39.2% (huge). Route easy volume to Haiku and hard work to Sonnet; at 20x per-token, one avoided Sonnet call pays for ~20 Haiku calls on the easy tier.
Is Haiku 5.5 good enough for coding?
For light code Q&A, review comments, and test triage — plausibly yes (Terminal-Bench 4.0 at 39.2% is a massive jump from Haiku 4.5's 0.0%). For agentic coding where the model must drive a terminal across many steps, the same-run gap to Sonnet is 31 points; that's a different job description.
Why does Haiku 5.5 get MORE expensive relative to Sonnet on big prompts?
Haiku tiers at 100K input tokens per request: $0.10/$0.50 below, $0.50/$2.50 above (a 5x step). Sonnet is flat. So the per-token gap shrinks from 20x to 4x as prompts grow — and Anthropic's new Haiku tokenizer uses slightly more tokens per task, widening real bills further.
Haiku vs Sonnet vs Opus — how do I choose?
Same-family ladder at flat-or-tiered list prices: Haiku $0.10/$0.50 (volume, routing, classification), Sonnet $2/$10 (best speed-intelligence balance, default effort high), Opus $4/$20 (maximum quality, the only one with Fast mode). Our calculator runs all three on your real token mix; the GPT-6.1 Sol pages cover the cross-vendor alternative at Sonnet's exact price.
Related
Sources & freshness
- Anthropic — Introducing Claude Haiku 5.5 (2026-10-07)https://www.anthropic.com/claude-haiku-5-5 · published 2026-10-07 · accessed 2026-10-09Tiered pricing ($0.10/$0.50 up to 100K, $0.50/$2.50 over) with Haiku 4.5 comparison ('90% lower up to 100K', 'on average ~75% less to run'); benchmark table includes GPT-6 Luna rows (GDPval-AA v2.1, AA-Briefcase v1.1, OSWorld 2.1, Terminal-Bench 4.0, FrontierCode 1.1, Chartography); 'fastest model to date' (at standard speed, slower than Opus in Fast Mode); tokenizer caveat 'uses slightly more tokens per task'.
- Anthropic — Claude models overview (model IDs, context window, max output, retirement)https://platform.claude.com/docs/en/docs/about-claude/models/overview · accessed 2026-10-09claude-opus-5-5: 1M context, 128K max output, training cutoff Jun 2026, retirement not sooner than 2027-09-22, cache reads '5% of input price', batch 50% off.
- Anthropic — Claude pricing (API pricing table incl. cache write/read, fast mode, batch)https://claude.com/pricing · accessed 2026-10-09Opus 5.5: $4/$20 per MTok, cache read $0.20, cache write (5m TTL) $5, batch 'save 50%', fast mode 2x on input/output.
- Anthropic — Claude pricing docs (cache TTL prices, fast/batch stacking rules)https://platform.claude.com/docs/en/about-claude/pricing · accessed 2026-10-09Opus 5.5 1h cache write $8/MTok; caching multipliers apply on top of fast-mode pricing and stack with the Batch discount; fast mode not available with Batch.
- Vals AI (independent) — Vals AI evaluation of Claude Haiku 5.5 (posted 2026-10-08)https://www.vals.ai/models/anthropic_claude-haiku-5-5 · published 2026-10-08 · accessed 2026-10-09Independent suite run at max effort, no refusal fallback: Vals Index 54.31% ±1.41 (#16/45); Terminal-Bench 4.0 35.35% (vs Anthropic's own 39.2% — different runs, both stored). No GPT-6 Luna datapoint.
- Artificial Analysis (independent) — Independent same-harness comparison: Claude Haiku 5.5 vs GPT-6 Luna (Intelligence Index v4.3.2)https://artificialanalysis.ai/models/comparisons/claude-haiku-5-5-vs-gpt-6-luna · accessed 2026-10-10Same-harness pair run. Corroborates Anthropic's own table (GDPval 1618/1432 vs 1620/1437). Cost-per-task $0.21 vs $0.07 at identical token prices (token verbosity).
Pricing data last verified 2026-10-09. OpenAI ships roughly weekly — re-verify before making spend commitments. See the methodology page for update cadence and stale-data flags.