Skip to content

What will Claude or GPT actually cost per task?

Per-token prices hide the real bill: reasoning verbosity, cache reads, and long-context cliffs. Every price and benchmark here is source-linked and dated, and we only compare scores measured in the same run.

Independent & unaffiliated with OpenAI/Anthropic · pricing verified 2026-10-09 · how we source data

Claude cost decisions

The comparisons where pricing structure — not marketing — decides: cliffs, cache, and independent same-harness quality data.

Compare by workload

Start from what you're building. Each card links to the comparison that covers that scenario.

Model cost calculator

Real API cost estimates, not list-price trivia: service tiers (batch/flex/fast/ultrafast), per-request long-context banding, cached reads and cache writes, monthly volume — for Claude Haiku, Sonnet, and Opus 5.5 and the GPT-6 family. Combinations the vendor hasn't priced show Unsupported, never a guess.

Calculate your workload cost →

Already have a usage export? Audit it against official rates · All tools

  • Coding agent on a mid-size repo: Agent loop reading ~12K tokens of repo context per request, ~1/3 of it re-served from prompt cache on later turns. Typical for Claude Code / Codex-style tools.
  • Long-document Q&A (crosses thresholds): One large document per request: 300K input tokens (100K cached from prior questions) — above OpenAI's 272K doubling threshold AND far above Haiku 5.5's 100K 5x cliff; only the flat-priced Claude models (Opus 5.5 and Sonnet 5.5) stay flat all the way.
  • High-volume support chat: Short exchanges where most of the system prompt + history is cached. A 'good enough quality' workload where GPT-6 Luna is often the right pick.

Model lineup & standard prices

Standard-tier list prices per 1M tokens; click through to the calculator for tier, cache, and long-context rates.

ModelInputCachedOutputReleased
Claude Haiku 5.5(Anthropic)For high-volume, latency-sensitive tasks such as classification, extraction, and routing (Anthropic docs).$0.10$0.01$0.502026-10-07
Claude Opus 5.5(Anthropic)$4$0.20$202026-09-22
Claude Sonnet 5.5(Anthropic)The best combination of speed and intelligence (Anthropic docs).$2$0.10$102026-09-28
GPT-6 Sol(OpenAI)$2$0.20$102026-09-22 *
GPT-6 Astra(OpenAI)Our most capable model for the most demanding work (OpenAI catalog tagline).$10$1$502026-09-03 *
GPT-6 Luna(OpenAI)Our most efficient model for focused, high-volume tasks (OpenAI catalog tagline).$0.10$0.01$0.502026-09-22 *
GPT-6.1 Sol(OpenAI)Near-Astra performance for complex work at a lower cost (OpenAI catalog tagline).$2$0.10$102026-09-29

USD per 1M tokens, standard tier. Anthropic · OpenAI · * release date from a secondary source · verified 2026-10-09 against OpenAI and Anthropic pricing.

Latest model updates

Source-linked official announcements only.

  1. Anthropic ships Claude Haiku 5.5 at GPT-6 Luna's exact price

    $0.10/$0.50 per MTok below 100K input tokens per request (5x above), 90% cheaper than Haiku 4.5, 'fastest model to date' at standard speed. Anthropic's own table shows it ahead of GPT-6 Luna on all six listed benchmarks; Vals AI independently scores it 54.31% overall.

    Source: anthropic.com

  2. OpenAI releases GPT-6.1 Sol (DevDay)

    Same $2/$10 list price as GPT-6 Sol, near-Astra evals, cached input halved to $0.10/MTok, Ultrafast tier priced. Available in ChatGPT Work and Codex at launch.

    Source: openai.com

  3. Anthropic answers with Claude Opus 5.5 the same day

    Opus 5.5 lands at $4/$20 per MTok — 20% below Opus 5 — with a 1M-token context window and 128K max output.

    Source: anthropic.com

How we compare models

Source-verified only

Every number links to the official page it came from, with the date we checked. Unverifiable claims don't render.

Shared-run rule

Benchmarks compare only when they share a documented run. Publishers mix their own evals with third-party and competitor-cited numbers (we record who published AND who ran each figure) — the same score can even differ across publications.

Tier-aware costs

Batch −50%, Fast 2×, Ultrafast 6×, long-context ≈2× above 272K input tokens per request, cache reads/writes billed where published.

Evidence-graded picks

Recommendations carry Measured / Inferred / Insufficient-evidence labels. 'We don't know yet' is a valid answer.

Read the full methodology →