Skip to content
GPT-6.1cross-vendor

GPT-6 Sol vs GPT-6.1 Sol: Upgrade Guide

Is GPT-6.1 Sol worth upgrading to?

The short answer

Headline numbers
At a glanceGPT-6.1 SolGPT-6 Sol
List priceinput / output per 1M tokens$2 / $10$2 / $10

Upgrade for almost everyone: GPT-6.1 Sol keeps the identical standard list price ($2 in / $10 out per MTok) while publishing better evals (DeepSWE +6.4 pp over 6 Sol's best at lower effort; factuality errors 11.4% → 7.7% at low effort) and halving the cached-input rate ($0.20 → $0.10). It also adds an Ultrafast tier 6 Sol never had. Nothing official forces the move — no deprecation timeline for gpt-6-sol has been announced — and behavioral compatibility beyond published evals is NOT officially documented, so run a regression on your own tasks before switching model IDs.

Choose GPT-6.1 Sol

Default for new and existing workloads: same list price, better published quality, half-price cache reads, ultrafast option.

Choose GPT-6 Sol

Keep temporarily if you need frozen behavior for reproducibility/compliance, or while your regression suite runs. Still billed normally; no deprecation announced as of 2026-10-09.

Key facts

  • List prices are identical at standard tier: $2 in / $10 out per MTok, short context; $4/$15 long. What changes: cached input halves from $0.20 to $0.10.
  • OpenAI announcement: 6.1 Sol beats 6 Sol's best DeepSWE v1.1 (68.8%) by 6.4 pp at a LOWER reasoning effort and cost; factuality error rate at low effort drops 11.4% → 7.7% (~32% fewer errors).
  • OSWorld 2.0: +7 pp over 6 Sol at max effort, at less than half the cost per task (OpenAI statement).
  • 6.1 Sol adds Ultrafast ($12/$60 short) — 6 Sol tops out at Fast ($4/$20).
  • API model ID changes: gpt-6-sol → gpt-6.1-sol (both remain billable; 6 Sol no longer appears in the catalog front page but is still on the pricing page at 2026-10-09).
  • NOT officially documented: response-format quirks, tool-calling behavioral parity, stop-token behavior, latency distributions. Treat all of that as UNKNOWN until you test.

Official pricing, every tier

USD per 1M tokens from vendor pricing pages. 'Not published' means exactly that — we don't extrapolate.

Official API pricing per million tokens, standard context band
Service tierGPT-6.1 SolGPT-6 Sol
Standarddefault on-demand serving

in $2 · cached $0.10 · out $10

cache write $2.5 · >272K in/req

in $2 · cached $0.20 · out $10

cache write $2.5 · >272K in/req

Batchasync jobs, 50% off

in $1 · cached $0.05 · out $5

cache write $1.25 · >272K in/req

in $1 · cached $0.10 · out $5

cache write $1.25 · >272K in/req

Flexslower/cheaper serving, 50% off

in $1 · cached $0.05 · out $5

cache write $1.25 · >272K in/req

in $1 · cached $0.10 · out $5

cache write $1.25 · >272K in/req

Fastpriority speed, 2x (ex-Priority)

in $4 · cached $0.20 · out $20

cache write $5 · >272K in/req

in $4 · cached $0.40 · out $20

cache write $5 · >272K in/req

Ultrafast6x standard price

in $12 · cached $0.60 · out $60

cache write $15 · >272K in/req

Not published
Long-context pricing per million tokens
Long-context bandGPT-6.1 SolGPT-6 Sol
Long-context ratesapplies per request above the model threshold

in $4 · cached $0.20 · out $15

cache write $5 · >272K in/req

in $4 · cached $0.40 · out $15

cache write $5 · >272K in/req

USD per 1M tokens. Every input token is billed in exactly one bucket — uncached, cached read, or cache write; a write replaces the input rate for those tokens (official rule), so write rates are per tier as listed.OpenAI long context = input > 272,000 tokens per single request. · OpenAI pricing (verified 2026-10-09)

Tier ≠ quality change

Batch/Flex/Fast/Ultrafast change serving speed and price, not the underlying model. Flex and Batch are half price; Fast is 2x; Ultrafast is 6x (GPT-6 family). On OpenAI these multipliers apply per tier as listed; Anthropic publishes Batch (-50%) and Fast (2x) for Claude Opus 5.5, and both stack with prompt-caching prices.

Run this comparison on your own numbers

Preloaded with a workload typical for this comparison. Change anything — results recompute locally.

Models (pick up to 4)

Estimated monthly cost per model
ModelPer requestMonthly
GPT-6.1 Sol cheapest$0.0674$1,348.00
GPT-6 Sol $0.0698$1,396.00

Monthly spread between the cheapest and priciest selected model: $48.00.

Breakdown: GPT-6.1 Sol · standard tier · standard context
  • Fresh input (uncached): 0 tok × $2/M$0.00
  • Cached input (reads): 24,000 tok × $0.1/M$0.0024
  • Cache writes: 6,000 tok × $2.5/M$0.015
  • Output (incl. reasoning tokens): 5,000 tok × $10/M$0.05
  • Per request$0.0674
  • 6,000 tokens billed at the cache-write rate instead of the input rate (mutually exclusive buckets — no double charge).

Rates: official pricing · verified 2026-10-09

  • Per-token arithmetic only: it does not model quality differences or cross-vendor output-token-count differences for the same task.
  • Reasoning tokens are billed as output tokens at the chosen model's output rate.
  • Long-context band is decided per request by total input tokens vs the provider threshold, not by monthly volume.
  • Cache accounting uses the official mutually-exclusive buckets: each input token is billed once, as uncached input, cached read, OR cache write — writes replace the input rate for those tokens, never add to it.
  • Claude rows bill the 5-minute-TTL cache write; the published 1-hour TTL ($8/MTok standard) is not modeled (no TTL input).
  • Rows flagged 'derived' are computed from a stated official rule (e.g. Anthropic batch -50%), not read verbatim from a price table.
  • Missing official prices render as Unsupported — we never apply a guessed multiplier (e.g., GPT-6 Sol has no Ultrafast tier).

What the evidence actually says

Independent and vendor-run evals, each with publisher, runner, and harness. Numbers from different runs are never merged.

Benchmark evidence log with harness attribution
BenchmarkScoreEffortHarnessEvidence
DeepSWEv1.1⚠ run comparability unknown
gpt-6.1-sol75.2 %highOpenAI announcement evalsVendor-reported
gpt-6-sol68.8 %best-reportedOpenAI announcement evals (quoted as prior best)Vendor-reported
AutomationBench1.0.6
gpt-6.1-sol+2.2 pp vs Claude Opus 5.5; +4.8 pp vs GPT-6 SolmediumOpenAI announcement (underlying runner not stated — AutomationBench is run/reported by Zapier per Anthropic's footnote)Vendor-reported
OSWorld2.0 offline set, v2026.08.08
gpt-6.1-solwithin 2.1 pp of GPT-6 Astra; +7 pp vs GPT-6 SolmaxOpenAI announcement evalsVendor-reported
Terminal-Bench-Science cost per task0.1
gpt-6.1-sol5.47 usd-per-taskunspecifiedOpenAI announcement evalsVendor-reported
Factuality error rateunspecified⚠ run comparability unknown
gpt-6.1-sol7.7 error-rate-%lowOpenAI announcement evalsVendor-reported
gpt-6-sol11.4 error-rate-%lowOpenAI announcement evals (quoted baseline)Vendor-reported
GDP.pdf (professional PDF Q&A)unspecified
gpt-6.1-solscores higher than Opus 5.5 with fallbacks at less than half the cost per task; near Astra at ~1/5 cost per taskmultiple settingsOpenAI announcement evalsVendor-reported
AA Intelligence Indexv4.3.2
gpt-6.1-sol52 index-pointsconfig 'Sol (Max)' per AAArtificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models)Third-party
Terminal-Bench4.0
gpt-6.1-sol56 %—Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models)Third-party
SciCodeunspecified
gpt-6.1-sol54 %—Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models)Third-party
GDPval-AAv2.1
gpt-6.1-sol1592 elo—Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models)Third-party
AA-Briefcasev1.1
gpt-6.1-sol1557 elo—Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models)Third-party
Humanity's Last Examunspecified
gpt-6.1-sol53 %—Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models)Third-party
AutomationBench-AAunspecified
gpt-6.1-sol65 %—Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models)Third-party
AA cost per Intelligence-Index taskv4.3.2
gpt-6.1-sol0.72 usd-per-task—Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models)Third-party

Same benchmark name, multiple rows — run comparability unknown

DeepSWE: gpt-6.1-sol = 75.2 (OpenAI announcement evals); gpt-6-sol = 68.8 (OpenAI announcement evals (quoted as prior best)). Same-harness comparison vs gpt-6-sol and gpt-6-astra only.

Factuality error rate: gpt-6.1-sol = 7.7 (OpenAI announcement evals); gpt-6-sol = 11.4 (OpenAI announcement evals (quoted baseline)). Whether these rows share a run is not documented in the sources; treat them as not comparable.

We never average or merge these. See methodology.

Sources: OpenAI — Introducing GPT-6.1 Sol (DevDay 2026) · Artificial Analysis (independent) — Independent same-harness comparison: GPT-6.1 Sol vs Claude Opus 5.5 (Intelligence Index v4.3.2)

Which model for which workload

Confidence labels: Measured = same-harness evidence · Inferred = reasoned from documented evidence · Insufficient evidence = no comparable public data.

Recommendations by workload scenario
WorkloadOur pickConfidenceWhy
Daily coding tasksGPT-6.1 SolMeasuredBetter published evals at identical list price; upgrade is close to free on quality-per-dollar.

Run your own regression first; behavior parity is not officially documented.

Budget-sensitive / high volumeGPT-6.1 SolMeasuredSame price, fewer errors at low effort → fewer retries; cache-heavy loops save directly via $0.10 cached reads.
Latency-sensitive productionGPT-6.1 SolPrice-derivedOnly 6.1 Sol offers Ultrafast in this pair; Fast tier pricing is identical ($4/$20).
Complex repository maintenanceGPT-6.1 SolMeasuredOSWorld +7 pp at max effort and DeepSWE +6.4 pp at lower effort point the same direction.

What officially changed (and what didn't)

Only confirmed rows come from vendor pages; 'Unknown' rows are gaps we refuse to fill with assumptions.

Officially confirmed vs unconfirmed API changes
AspectBefore (gpt-6-sol)After (gpt-6.1-sol)Status
Model IDgpt-6-solgpt-6.1-solConfirmed
Standard input / output price (short ctx)$2.00 / $10.00$2.00 / $10.00Confirmed
Cached input price (short ctx)$0.20$0.10Confirmed
Ultrafast tiernot offered$12 / $60 (short ctx)Confirmed
Reasoning effortslow / medium / high / xhigh / maxlow / medium / high / xhigh / maxConfirmed
Prompt/tool behavior parity—not officially documentedUnknown
Deprecation of gpt-6-sol—no timeline announced as of 2026-10-09Unknown

When staying on the old model is reasonable

  • You need frozen, reproducible outputs for compliance or benchmarking (no vendor guarantees behavioral freezes — freezing means 'don't change anything', including the model ID).
  • Your regression suite hasn't run yet — that's a reason to wait days, not months: the upgrade is same-price on standard tier.
  • You depend on a third-party gateway that hasn't enabled gpt-6.1-sol (check with the provider; don't assume).

GPT-6 Sol → GPT-6.1 Sol migration checklist

Behavioral compatibility is not officially documented, so the checklist is: freeze a golden set, swap the ID in staging, measure, then roll out gradually. Progress is saved in your browser only and can be exported.

GPT-6 Sol → GPT-6.1 Sol migration checklist

0/8 complete

Behavioral compatibility is not officially documented, so the checklist is: freeze a golden set, swap the ID in staging, measure, then roll out gradually. Progress is saved in your browser only and can be exported.

State is stored in your browser's localStorage only — nothing is uploaded, and the export file is generated locally.

Frequently asked questions

Is gpt-6-sol being deprecated?

No timeline has been announced as of 2026-10-09. It was removed from the model catalog front page but remains on the official pricing page and fully billable. We monitor the OpenAI changelog and update this page when that changes.

Will my prompts behave identically after switching the model ID?

That is not officially documented. Published evals show quality improvements, which by definition means behavior changed. Freeze a golden set of your real tasks, run both IDs, and diff the outputs — the checklist above walks you through it.

Does upgrading change my bill?

Standard list prices are identical ($2/$10). Two real effects: cached input halves to $0.10/MTok (cache-heavy agent loops get cheaper), and the optional Ultrafast tier ($12/$60) adds a new way to spend — only if you opt in. If 6.1 Sol reaches your quality bar at a lower reasoning effort, per-task cost can drop further.

What did OpenAI officially confirm as different?

From the GPT-6.1 Sol announcement: better DeepSWE v1.1 (+6.4 pp over 6 Sol's best, at lower effort), factuality error rate 11.4% → 7.7% at low effort, OSWorld +7 pp at max, cached input 50% cheaper, Ultrafast tier availability. Everything else — tool-calling parity, latency, formatting — is unconfirmed.

Sources & freshness

  • OpenAI — OpenAI API model catalog (model IDs, context windows, reasoning efforts, capabilities)https://developers.openai.com/api/docs/models · accessed 2026-10-09GPT-6 Sol is NOT listed in the catalog front page at access time (still priced on the pricing page). Knowledge cutoffs: Apr 30 2026 (Astra/6.1 Sol), May 18 2026 (Luna).
  • OpenAI — OpenAI API pricing (GPT-6 family tier matrix: standard/batch/flex/fast/ultrafast, short/long context, cache)https://developers.openai.com/api/docs/pricing · accessed 2026-10-09Long context defined as >272K input tokens per request (2x input and cache rates, 1.5x output for the whole request). 'Priority processing' renamed 'Fast mode' on 2026-07-30. Flex and Batch are half of Standard for the GPT-6 family, Fast is 2x, Ultrafast is 6x. Cache writes cost 1.25x the uncached input rate, reads 0.1x (0.05x on GPT-6.1 Sol). Cache-write rates are captured on every published tier row, not Standard only.
  • OpenAI — Introducing GPT-6.1 Sol (DevDay 2026)https://openai.com/index/introducing-gpt-6-1-sol · published 2026-09-29 · accessed 2026-10-09Self-reported evals (OpenAI harness). Includes per-task cost quotes (Terminal-Bench-Science). Announces GPT-6.1 Sol Ultrafast 'coming soon' in Codex. Availability: ChatGPT Work + Codex, not yet in Chat.
  • Artificial Analysis (independent) — Independent same-harness comparison: GPT-6.1 Sol vs Claude Opus 5.5 (Intelligence Index v4.3.2)https://artificialanalysis.ai/models/comparisons/gpt-6-1-sol-vs-claude-opus-5-5 · accessed 2026-10-10Independent same-harness evaluation of both models. Cost-per-task = weighted average per Intelligence Index task; reflects measured token verbosity. Effort configs as labeled by AA ('Sol (Max)' vs 'Opus 5.5 (Max, Default Fallback)').

Pricing data last verified 2026-10-09. OpenAI ships roughly weekly — re-verify before making spend commitments. See the methodology page for update cadence and stale-data flags.