Skip to content

How we compare models

This page explains what our numbers mean, where each one comes from, and the limits of the whole exercise. It exists so you can audit us — not to bury caveats.

1. Data sources

We use two tiers of sources. Official: vendor pricing pages, model catalogs, and announcement posts (OpenAI developers docs, openai.com announcements, claude.com pricing, Anthropic docs and announcements) — every price and benchmark on this site links to one, with the date we verified it (latest: 2026-10-09). Secondary: our own pre-launch research notes, used only for facts like release dates that the fetched pages don't restate; those carry an explicit "unconfirmed" badge everywhere they appear. We do not scrape or republish third-party benchmark datasets; when we cite a vendor's table, we link to it rather than copying the whole dataset.

2. Official data vs published benchmarks

Pricing and model metadata (IDs, context windows, tiers) are official facts — we verify them against vendor pages and date-stamp each row. Benchmarks are different: nearly everything published right now is vendor-published but not necessarily vendor-run. Attribution footnotes matter: Anthropic's Opus 5.5 table credits AutomationBench to Zapier (who ran and reported it, with no fallback models), and footnotes several GPT-6 Astra figures as "as reported by OpenAI"; OpenAI's GPT-6.1 Sol footnote says competitor evaluations were "taken from publicly available reports". Worse, the same OpenAI-published Astra score differs between the two announcements (Terminal-Bench-Science: 68.1% vs 64.6%), with the run/version differences undocumented.

The shared-run rule

Numbers compare only when they share a documented run. Publisher ≠ runner: we record both on every row. When two publications disagree about the same figure, we store both and mark the cause UNKNOWN — we never average, merge, or invent a reconciliation.

3. Reasoning effort levels

OpenAI GPT-6 family models accept reasoning efforts low/medium/high/xhigh/max (Luna also none). Higher effort trades latency and output tokens (which are billed as output) for quality. Anthropic's Claude Opus 5.5 uses "adaptive thinking" that is always on, with effort tiers Low→Max. Benchmark rows on this site always state the effort level used, because a score at max effort vs default effort is not a fair fight — and cost-per-task claims depend on it (OpenAI notes GPT-6.1 Sol hits its DeepSWE score at a lower effort than GPT-6 Sol's best).

4. Service tiers

OpenAI: Batch and Flex at 50% of Standard (async / relaxed serving), Fast at 2× (priority serving; renamed from "Priority" on 2026-07-30), Ultrafast at 6× (not offered on GPT-6 Sol or Luna). Anthropic: Batch at −50%, Fast mode at 2× on input/output. Tiers change serving speed and price, not the model. Cache-write prices are published only for some tiers; the calculator marks the rest Unsupported instead of guessing.

5. Long context and cache billing

OpenAI doubles GPT-6 family rates when a single request carries more than 272,000 input tokens — judged per request, never accumulated over a month. Anthropic's context pricing is model-dependent: Claude Opus 5.5 has no surcharge (flat to its 1M window), while Claude Haiku 5.5 steps up 5x above 100K input tokens per request — a much lower cliff than OpenAI's, at a much higher multiplier. Prompt caches on both platforms use mutually exclusive token buckets: every input token is billed exactly once — as uncached input, a cached read, or a cache write — so a write replaces the input rate for those tokens instead of adding to it. OpenAI publishes write rates per tier (e.g. GPT-6.1 Sol $2.50 standard short, $1.25 batch, $5 fast, $15 ultrafast); Anthropic prices 5-minute writes at $5/MTok for Opus 5.5 and 1-hour writes at $8/MTok, with reads at 5% of input — and these multipliers stack with the Batch discount and Fast-mode pricing. Our calculator models the 5m TTL (no TTL input yet); OpenAI requires a 1,024-token minimum cacheable prefix, and cache reads refresh the entry for free on both platforms. Cache economics dominate long-lived agent loops — model them, don't ignore them.

6. Recommendation rules & evidence confidence

  • Measured — a pick backed by same-harness evidence (typically the vendor's own table comparing its models, or official pricing arithmetic).
  • Inferred — reasoned from documented evidence (e.g. vendor case studies about long-horizon coding), not directly measured on one harness.
  • Insufficient evidence — no comparable public data exists. We say this out loud and hand you a bake-off protocol instead of a fake verdict.

7. Update cadence & example workloads

OpenAI shipped five notable GPT-6-family releases between Sep 3 and Oct 7, 2026 — roughly weekly. We re-verify pricing on every page edit and treat any row older than ~14 days as due for a re-check. Release dates and other facts from secondary research stay flagged until re-confirmed. The calculator's example workloads (coding agent, long-doc QA, support chat, nightly batch, cache-heavy fleet) are deliberately varied so tier/cache/band edge cases are one click to explore.

8. Commercial model & editorial independence

This site currently has no ads, no affiliate links, and no sponsored content. We don't host API keys, don't take payments, and the "Model Audit" early-access panel describes a future product without pretending it exists. If monetization is ever added (affiliate links, paid audits), vendor relationships will be disclosed on this page and will never change a verdict, a price, or an evidence label. We are independent and unaffiliated with OpenAI and Anthropic.

9. Privacy

The calculator runs entirely in your browser; token counts and exports never leave your device. We measure aggregate product analytics (page views, tool usage counts) with no PII, no API keys, and no bill data — internal/QA traffic is filtered. There is no account system and no email capture today.