Model cost calculator
Per-token list prices are public — your actual bill isn't. This calculator models what actually moves it: service tier, cached-input share, cache writes, per-request long-context banding, and monthly volume, across Claude Haiku, Sonnet, and Opus 5.5 and the GPT-6 family.
Examples are explained on the methodology page.
| Model | Per request | Monthly |
|---|---|---|
| GPT-6.1 Sol cheapest | $0.0414 | $1,656.00 |
| Claude Opus 5.5 | $0.0828 | $3,312.00 |
Monthly spread between the cheapest and priciest selected model: $1,656.00.
Breakdown: GPT-6.1 Sol · standard tier · standard context
- Fresh input (uncached): 8,000 tok × $2/M$0.016
- Cached input (reads): 4,000 tok × $0.1/M$0.0004
- Cache writes: 0 tok × $2.5/M$0.00
- Output (incl. reasoning tokens): 2,500 tok × $10/M$0.025
- Per request$0.0414
Rates: official pricing · verified 2026-10-09
- Per-token arithmetic only: it does not model quality differences or cross-vendor output-token-count differences for the same task.
- Reasoning tokens are billed as output tokens at the chosen model's output rate.
- Long-context band is decided per request by total input tokens vs the provider threshold, not by monthly volume.
- Cache accounting uses the official mutually-exclusive buckets: each input token is billed once, as uncached input, cached read, OR cache write — writes replace the input rate for those tokens, never add to it.
- Claude rows bill the 5-minute-TTL cache write; the published 1-hour TTL ($8/MTok standard) is not modeled (no TTL input).
- Rows flagged 'derived' are computed from a stated official rule (e.g. Anthropic batch -50%), not read verbatim from a price table.
- Missing official prices render as Unsupported — we never apply a guessed multiplier (e.g., GPT-6 Sol has no Ultrafast tier).
How the math works
The formula
cost = fresh_in × in_rate + cached_in × cached_rate + cache_writes × write_rate + out × out_rate, all ÷ 1,000,000, × requests/month
Rates come from official pricing pages per (model, tier, context band). Rows marked "derived" apply a stated official rule (e.g. Anthropic batch −50%).
What this estimate is NOT
It is arithmetic, not a forecast of your invoice. Quality changes, retry rates, and cross-vendor output-token differences (models don't write the same number of tokens for one task) are outside per-token math. Where a vendor hasn't published a rate (e.g. cache under some tiers), the combination is Unsupported — we never extrapolate a multiplier.
Full details, example workloads, and data-source policy: methodology & data sources.