Claude Haiku 5.5 vs GPT-6 Luna
They cost the same on paper — which small model is actually cheaper and better for your workload?
The short answer
| At a glance | Claude Haiku 5.5 | GPT-6 Luna |
|---|---|---|
| List priceinput / output per 1M tokens | $0.10 / $0.50 | $0.10 / $0.50 |
| Cost per taskArtificial Analysis, measured | $0.21 | $0.07 (better) |
| Intelligence IndexAA v4.3.2, higher is better | 43 (better) | 38 |
Identical list price below the thresholds ($0.10/$0.50, $0.01 cache reads) — and now two independent evidence sources. Artificial Analysis (same harness, Index v4.3.2): Haiku 5.5 leads overall quality (Index 43 vs 38; Terminal-Bench 33% vs 13%) — corroborating Anthropic's own table (GDPval-AA 1620 vs 1437 there; 1618 vs 1432 at AA — two sources, same story). But the pair is task-shaped, not a blanket lead: Luna WINS AutomationBench-AA 53% vs 35%. And per-task costs diverge at identical token prices: Haiku $0.21 vs Luna $0.07 per AA task, because Haiku thinks ~3x more tokens (162k vs 50k output). Conditional rule: under ~100K inputs and quality-first → Haiku; any request over 100K (5x cliff), cost-per-task volume, or fast-first-token (Luna 102s vs 296s TTFT) → Luna; throughput favors Haiku (240 vs 137 tok/s).
Choose Claude Haiku 5.5
Sub-100K prompts where quality-first: Index 43 vs 38, Terminal-Bench 33% vs 13% (same-harness AA; Anthropic's table agrees), highest throughput (240 vs 137 tok/s).
Choose GPT-6 Luna
Anything crossing 100K inputs (Haiku's 5x tier), task-cost-sensitive volume ($0.07 vs $0.21 per AA task), or latency-to-first-token (102s vs 296s); also Luna's batch rates are explicit while Haiku's are derived.
Either — decide by evidence you generate
Sub-100K price-identical traffic with mixed tasks — note Luna wins AutomationBench-AA 53% vs 35%, so 'Haiku is the smart one' is not a blanket truth.
Key facts
- Identical list prices below the thresholds: $0.10 in / $0.50 out / $0.01 cached read per MTok — Haiku for prompts ≤100K, Luna for ≤272K (verified 2026-10-09/10).
- INDEPENDENT same-harness (AA Index v4.3.2): Haiku 5.5 Index 43 vs Luna 38; Terminal-Bench 4.0 33% vs 13%; GDPval-AA 1618 vs 1432; AA-Briefcase 1577 vs 1336 — corroborating Anthropic's own table (1620/1437, 1578/1336) from a second source.
- The blanket lead has exceptions: Luna WINS AutomationBench-AA 53% vs 35% — route business-automation-style work accordingly.
- Per-task cost at identical token prices (AA): Haiku $0.21 vs Luna $0.07 — Haiku emits ~3x the output tokens per task (162k vs 50k; 129k vs 39k reasoning). Per-token parity ≠ per-task parity.
- The cliffs: Haiku multiplies by 5 above 100K input tokens/request ($0.50/$2.50); Luna doubles above 272K ($0.20/$0.75). A 120K-in/1.5K-out request: $0.0548 on Haiku vs $0.0110 on Luna.
- Latency split (AA): Haiku 240 vs 137 tok/s throughput; Luna reaches first token 102s vs 296s.
- Batch: Luna explicit $0.05/$0.25; Haiku derived −50% ($0.05/$0.25 below 100K, $0.25/$1.25 above). Luna offers Fast (2x); Haiku lists no fast/flex tier.
- Vs Haiku 4.5 (official): 90% cheaper input below 100K, ~75% less on average; tokenizer uses slightly more tokens per task.
Official pricing, every tier
USD per 1M tokens from vendor pricing pages. 'Not published' means exactly that — we don't extrapolate.
| Service tier | Claude Haiku 5.5 | GPT-6 Luna |
|---|---|---|
| Standarddefault on-demand serving | in $0.10 · cached $0.01 · out $0.50 cache write $0.125 · >100K in/req | in $0.10 · cached $0.01 · out $0.50 cache write $0.125 · >272K in/req |
| Batchasync jobs, 50% off | in $0.05 · cached $0.005 · out $0.25 cache write $0.0625 · derived · >100K in/req | in $0.05 · cached $0.005 · out $0.25 cache write $0.0625 · >272K in/req |
| Flexslower/cheaper serving, 50% off | Not published | in $0.05 · cached $0.005 · out $0.25 cache write $0.0625 · >272K in/req |
| Fastpriority speed, 2x (ex-Priority) | Not published | in $0.20 · cached $0.02 · out $1 cache write $0.25 · >272K in/req |
| Long-context band | Claude Haiku 5.5 | GPT-6 Luna |
|---|---|---|
| Long-context ratesapplies per request above the model threshold | in $0.50 · cached $0.05 · out $2.5 cache write $0.625 · >100K in/req | in $0.20 · cached $0.02 · out $0.75 cache write $0.25 · >272K in/req |
USD per 1M tokens. Every input token is billed in exactly one bucket — uncached, cached read, or cache write; a write replaces the input rate for those tokens (official rule), so write rates are per tier as listed.OpenAI long context = input > 272,000 tokens per single request. · OpenAI pricing (verified 2026-10-09) · Anthropic rows bill the 5-minute cache-write TTL; the 1-hour TTL ($8/MTok standard) is also published, and cache multipliers stack with the Batch discount and Fast-mode pricing. · Anthropic pricing (verified 2026-10-09)
Tier ≠ quality change
Run this comparison on your own numbers
Preloaded with a workload typical for this comparison. Change anything — results recompute locally.
| Model | Per request | Monthly |
|---|---|---|
| Claude Haiku 5.5 long ctx | $0.05475 | $2,737.50 |
| GPT-6 Luna cheapest | $0.01095 | $547.50 |
Monthly spread between the cheapest and priciest selected model: $2,190.00.
Breakdown: Claude Haiku 5.5 · standard tier · long context
- Fresh input (uncached): 100,000 tok × $0.5/M$0.05
- Cached input (reads): 20,000 tok × $0.05/M$0.001
- Cache writes: 0 tok × $0.625/M$0.00
- Output (incl. reasoning tokens): 1,500 tok × $2.5/M$0.00375
- Per request$0.05475
- Long-context band applies: input of 120,000 tokens exceeds the 100000-per-request threshold.
Rates: official pricing · verified 2026-10-09
- Per-token arithmetic only: it does not model quality differences or cross-vendor output-token-count differences for the same task.
- Reasoning tokens are billed as output tokens at the chosen model's output rate.
- Long-context band is decided per request by total input tokens vs the provider threshold, not by monthly volume.
- Cache accounting uses the official mutually-exclusive buckets: each input token is billed once, as uncached input, cached read, OR cache write — writes replace the input rate for those tokens, never add to it.
- Claude rows bill the 5-minute-TTL cache write; the published 1-hour TTL ($8/MTok standard) is not modeled (no TTL input).
- Rows flagged 'derived' are computed from a stated official rule (e.g. Anthropic batch -50%), not read verbatim from a price table.
- Missing official prices render as Unsupported — we never apply a guessed multiplier (e.g., GPT-6 Sol has no Ultrafast tier).
What the evidence actually says
Independent and vendor-run evals, each with publisher, runner, and harness. Numbers from different runs are never merged.
| Benchmark | Score | Effort | Harness | Evidence |
|---|---|---|---|---|
| GDPval-AAv2.1⚠ run comparability unknown | ||||
| claude-haiku-5-5 | 1620 elo | unspecified | Anthropic Haiku 5.5 announcement evals (GDPval-AA benchmark by Artificial Analysis) | Vendor-reported |
| gpt-6-luna | 1437 elo | unspecified | Anthropic Haiku 5.5 announcement evals (competitor row) | Vendor-reported |
| claude-haiku-5-5 | 1618 elo | — | Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models) | Third-party |
| gpt-6-luna | 1432 elo | — | Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models) | Third-party |
| AA-Briefcasev1.1⚠ run comparability unknown | ||||
| claude-haiku-5-5 | 1578 elo | unspecified | Anthropic Haiku 5.5 announcement evals | Vendor-reported |
| gpt-6-luna | 1336 elo | unspecified | Anthropic Haiku 5.5 announcement evals (competitor row) | Vendor-reported |
| claude-haiku-5-5 | 1577 elo | — | Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models) | Third-party |
| gpt-6-luna | 1336 elo | — | Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models) | Third-party |
| OSWorld2.1 offline subset | ||||
| claude-haiku-5-5 | 72.4 % | unspecified | Anthropic Haiku 5.5 announcement evals | Vendor-reported |
| gpt-6-luna | 48.9 % | unspecified | Anthropic Haiku 5.5 announcement evals (competitor row) | Vendor-reported |
| Terminal-Bench4.0⚠ run comparability unknown | ||||
| claude-haiku-5-5 | 39.2 % | unspecified | Anthropic Haiku 5.5 announcement evals | Vendor-reported |
| gpt-6-luna | 16.4 % | unspecified | Anthropic Haiku 5.5 announcement evals (competitor row) | Vendor-reported |
| claude-haiku-5-5 | 35.35 % | max | Vals AI (independent) | Third-party |
| claude-haiku-5-5 | 33 % | — | Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models) | Third-party |
| gpt-6-luna | 13 % | — | Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models) | Third-party |
| FrontierCode1.1 (Main) | ||||
| claude-haiku-5-5 | 46.4 % | unspecified | Anthropic Haiku 5.5 announcement evals | Vendor-reported |
| gpt-6-luna | 42.4 % | xhigh | Anthropic Haiku 5.5 announcement evals (competitor row) | Vendor-reported |
| Chartographyunspecified (no tools) | ||||
| claude-haiku-5-5 | 46.4 % | unspecified | Anthropic Haiku 5.5 announcement evals | Vendor-reported |
| gpt-6-luna | 29.1 % | unspecified | Anthropic Haiku 5.5 announcement evals (competitor row) | Vendor-reported |
| Vals Index (overall)unspecified | ||||
| claude-haiku-5-5 | 54.31 % | max | Vals AI (independent, no refusal fallback) | Third-party |
| AA Intelligence Indexv4.3.2 | ||||
| claude-haiku-5-5 | 43 index-points | — | Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models) | Third-party |
| gpt-6-luna | 38 index-points | — | Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models) | Third-party |
| AutomationBench-AAunspecified | ||||
| claude-haiku-5-5 | 35 % | — | Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models) | Third-party |
| gpt-6-luna | 53 % | — | Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models) | Third-party |
| AA cost per Intelligence-Index taskv4.3.2 | ||||
| claude-haiku-5-5 | 0.21 usd-per-task | — | Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models) | Third-party |
| gpt-6-luna | 0.07 usd-per-task | — | Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models) | Third-party |
Same benchmark name, multiple rows — run comparability unknown
GDPval-AA: claude-haiku-5-5 = 1620 (Anthropic Haiku 5.5 announcement evals (GDPval-AA benchmark by Artificial Analysis)); gpt-6-luna = 1437 (Anthropic Haiku 5.5 announcement evals (competitor row)); claude-haiku-5-5 = 1618 (Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models)); gpt-6-luna = 1432 (Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models)). Same table contains the GPT-6 Luna row (1437) — a same-table, same-run comparison of exactly the pair our haiku-vs-luna page covers. Column order in the source table is Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | Sonnet 5.5; 735 belongs to Haiku 4.5 and MUST NOT be read as Luna's value (this exact column shift was a published error, corrected 2026-10-10). Different table from the Opus 5.5 announcement's GDPval (Opus 1846 / Astra 1542): never merge across tables.
AA-Briefcase: claude-haiku-5-5 = 1578 (Anthropic Haiku 5.5 announcement evals); gpt-6-luna = 1336 (Anthropic Haiku 5.5 announcement evals (competitor row)); claude-haiku-5-5 = 1577 (Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models)); gpt-6-luna = 1336 (Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models)). Whether these rows share a run is not documented in the sources; treat them as not comparable.
Terminal-Bench: claude-haiku-5-5 = 39.2 (Anthropic Haiku 5.5 announcement evals); gpt-6-luna = 16.4 (Anthropic Haiku 5.5 announcement evals (competitor row)); claude-haiku-5-5 = 35.35 (Vals AI (independent)); claude-haiku-5-5 = 33 (Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models)); gpt-6-luna = 13 (Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models)). Vals AI independently measured Haiku 5.5 at 35.35% on Terminal-Bench 4.0 (see terminalbench4-haiku55-vals) — vendor vs independent runs differ by ~4 pp; both stored, neither merged.
We never average or merge these. See methodology.Sources: Anthropic — Introducing Claude Haiku 5.5 (2026-10-07) · Vals AI (independent) — Vals AI evaluation of Claude Haiku 5.5 (posted 2026-10-08) · Artificial Analysis (independent) — Independent same-harness comparison: Claude Haiku 5.5 vs GPT-6 Luna (Intelligence Index v4.3.2)
Which model for which workload
Confidence labels: Measured = same-harness evidence · Inferred = reasoned from documented evidence · Insufficient evidence = no comparable public data.
| Workload | Our pick | Confidence | Why |
|---|---|---|---|
| Daily coding tasks | Claude Haiku 5.5 | Measured | Terminal-Bench 4.0 33% vs 13% same-harness (AA); Anthropic's table (39.2% vs 16.4%) agrees in direction. |
| Agent workflows | Decide by bake-off | Measured | Split by task type: Anthropic's OSWorld 72.4% vs 48.9% favors Haiku, but AA's AutomationBench-AA 53% vs 35% favors Luna — genuinely task-shaped; run your own mix before committing a fleet. Two independent-ish sources disagree by task family — that's the honest answer. |
| Long-context tasks | GPT-6 Luna | Price-derived | Cliff arithmetic: >100K puts Haiku at $0.50/$2.50 while Luna holds $0.10/$0.50 to 272K — Luna ~5x cheaper at 120K input. |
| Budget-sensitive / high volume | GPT-6 Luna | Measured | $0.07 vs $0.21 per AA-measured task at identical token prices; below 100K list prices tie, so the task-cost gap is pure token verbosity. |
| Quality-first | Claude Haiku 5.5 | Measured | AA Index 43 vs 38 plus Terminal-Bench lead; Vals AI independently rates Haiku 54.31% overall. |
| Latency-sensitive production | GPT-6 Luna | Measured | First token 102s vs 296s (AA) — for streaming/chat feel; Haiku wins raw throughput 240 vs 137 tok/s. |
| Offline batch jobs | GPT-6 Luna | Price-derived | Explicit batch rates + per-task cost lead; above 100K Luna's batch ($0.10/$0.375) beats Haiku's derived ($0.25/$1.25). |
| Complex repository maintenance | Decide by bake-off | Inferred | Small models, big repos — if forced, Haiku's TB4 edge applies; the real answer is the mid/frontier tier. |
Frequently asked questions
What does Claude Haiku 5.5 cost?
$0.10 per million input tokens and $0.50 per million output tokens for requests up to 100K input tokens; above that, $0.50/$2.50. Cache reads $0.01 ($0.05 over 100K); cache writes are $0.125 for the 5-minute TTL and $0.20 for the 1-hour TTL ($0.625 / $1.00 over 100K), Batch −50%. That makes it 90% cheaper than Haiku 4.5 below 100K and ~75% cheaper on average, per Anthropic's announcement.
Is Haiku 5.5 the same price as GPT-6 Luna?
Yes — exactly, up to a point: both are $0.10/$0.50 per MTok with $0.01 cache reads. The divergence is where they step up: Haiku multiplies by 5 above 100K input tokens per request; Luna doubles above 272K. Any request between 100K and 272K tokens is cheaper on Luna — and AA's per-task measurements ($0.21 vs $0.07) show the gap shows up even below the cliff via token verbosity.
Which is better, Haiku 5.5 or GPT-6 Luna?
The only same-run public comparison is Anthropic's own announcement table, where Haiku 5.5 leads Luna on all six listed benchmarks — by the widest margins on computer use (OSWorld 72.4% vs 48.9%) and agentic coding (Terminal-Bench 4.0 39.2% vs 16.4%), and by the narrowest on professional knowledge work (GDPval-AA 1620 vs 1437, AA-Briefcase 1578 vs 1336). Independent Vals AI data confirms Haiku is strong (54.31% overall index) but publishes no Luna number, so the head-to-head verdict rests on one vendor's table. Run both on 30–50 of your own tasks before committing.
What happens above 100K tokens on Haiku 5.5?
Every rate steps up 5x: input $0.50, output $2.50, cache reads $0.05, cache writes $0.625 per MTok. The tier is judged per request, never monthly. If your prompts regularly cross 100K, either switch those workloads to Luna (holds $0.10/$0.50 to 272K) or restructure prompts to stay under the cliff.
Does Haiku 5.5 have a fast mode?
No fast or flex tier is listed on Anthropic's pricing page for Haiku 5.5 (standard and batch only). GPT-6 Luna has Fast at 2x ($0.20/$1.00). Anthropic positions Haiku 5.5 as its fastest model at standard speed — with its own footnote that this is still slower than Opus 5.5 in Fast mode.
Is Haiku 5.5 cheaper than Haiku 4.5?
Yes, dramatically: $0.10 vs $1.00 input below 100K (90% lower per Anthropic), $0.50 vs $1.00 above (50% lower), and 'on average around 75% less to run'. One caveat from Anthropic itself: the new tokenizer uses slightly more tokens per task, so real savings are somewhat below the headline percentages.
Related
Sources & freshness
- Anthropic — Introducing Claude Haiku 5.5 (2026-10-07)https://www.anthropic.com/claude-haiku-5-5 · published 2026-10-07 · accessed 2026-10-09Tiered pricing ($0.10/$0.50 up to 100K, $0.50/$2.50 over) with Haiku 4.5 comparison ('90% lower up to 100K', 'on average ~75% less to run'); benchmark table includes GPT-6 Luna rows (GDPval-AA v2.1, AA-Briefcase v1.1, OSWorld 2.1, Terminal-Bench 4.0, FrontierCode 1.1, Chartography); 'fastest model to date' (at standard speed, slower than Opus in Fast Mode); tokenizer caveat 'uses slightly more tokens per task'.
- OpenAI — OpenAI API model catalog (model IDs, context windows, reasoning efforts, capabilities)https://developers.openai.com/api/docs/models · accessed 2026-10-09GPT-6 Sol is NOT listed in the catalog front page at access time (still priced on the pricing page). Knowledge cutoffs: Apr 30 2026 (Astra/6.1 Sol), May 18 2026 (Luna).
- OpenAI — OpenAI API pricing (GPT-6 family tier matrix: standard/batch/flex/fast/ultrafast, short/long context, cache)https://developers.openai.com/api/docs/pricing · accessed 2026-10-09Long context defined as >272K input tokens per request (2x input and cache rates, 1.5x output for the whole request). 'Priority processing' renamed 'Fast mode' on 2026-07-30. Flex and Batch are half of Standard for the GPT-6 family, Fast is 2x, Ultrafast is 6x. Cache writes cost 1.25x the uncached input rate, reads 0.1x (0.05x on GPT-6.1 Sol). Cache-write rates are captured on every published tier row, not Standard only.
- Anthropic — Claude pricing (API pricing table incl. cache write/read, fast mode, batch)https://claude.com/pricing · accessed 2026-10-09Opus 5.5: $4/$20 per MTok, cache read $0.20, cache write (5m TTL) $5, batch 'save 50%', fast mode 2x on input/output.
- Anthropic — Claude pricing docs (cache TTL prices, fast/batch stacking rules)https://platform.claude.com/docs/en/about-claude/pricing · accessed 2026-10-09Opus 5.5 1h cache write $8/MTok; caching multipliers apply on top of fast-mode pricing and stack with the Batch discount; fast mode not available with Batch.
- Vals AI (independent) — Vals AI evaluation of Claude Haiku 5.5 (posted 2026-10-08)https://www.vals.ai/models/anthropic_claude-haiku-5-5 · published 2026-10-08 · accessed 2026-10-09Independent suite run at max effort, no refusal fallback: Vals Index 54.31% ±1.41 (#16/45); Terminal-Bench 4.0 35.35% (vs Anthropic's own 39.2% — different runs, both stored). No GPT-6 Luna datapoint.
- Artificial Analysis (independent) — Independent same-harness comparison: Claude Haiku 5.5 vs GPT-6 Luna (Intelligence Index v4.3.2)https://artificialanalysis.ai/models/comparisons/claude-haiku-5-5-vs-gpt-6-luna · accessed 2026-10-10Same-harness pair run. Corroborates Anthropic's own table (GDPval 1618/1432 vs 1620/1437). Cost-per-task $0.21 vs $0.07 at identical token prices (token verbosity).
Pricing data last verified 2026-10-09. OpenAI ships roughly weekly — re-verify before making spend commitments. See the methodology page for update cadence and stale-data flags.