GPT-6 Sol vs GPT-6.1 Sol: Upgrade Guide
Is GPT-6.1 Sol worth upgrading to?
The short answer
| At a glance | GPT-6.1 Sol | GPT-6 Sol |
|---|---|---|
| List priceinput / output per 1M tokens | $2 / $10 | $2 / $10 |
Upgrade for almost everyone: GPT-6.1 Sol keeps the identical standard list price ($2 in / $10 out per MTok) while publishing better evals (DeepSWE +6.4 pp over 6 Sol's best at lower effort; factuality errors 11.4% → 7.7% at low effort) and halving the cached-input rate ($0.20 → $0.10). It also adds an Ultrafast tier 6 Sol never had. Nothing official forces the move — no deprecation timeline for gpt-6-sol has been announced — and behavioral compatibility beyond published evals is NOT officially documented, so run a regression on your own tasks before switching model IDs.
Choose GPT-6.1 Sol
Default for new and existing workloads: same list price, better published quality, half-price cache reads, ultrafast option.
Choose GPT-6 Sol
Keep temporarily if you need frozen behavior for reproducibility/compliance, or while your regression suite runs. Still billed normally; no deprecation announced as of 2026-10-09.
Key facts
- List prices are identical at standard tier: $2 in / $10 out per MTok, short context; $4/$15 long. What changes: cached input halves from $0.20 to $0.10.
- OpenAI announcement: 6.1 Sol beats 6 Sol's best DeepSWE v1.1 (68.8%) by 6.4 pp at a LOWER reasoning effort and cost; factuality error rate at low effort drops 11.4% → 7.7% (~32% fewer errors).
- OSWorld 2.0: +7 pp over 6 Sol at max effort, at less than half the cost per task (OpenAI statement).
- 6.1 Sol adds Ultrafast ($12/$60 short) — 6 Sol tops out at Fast ($4/$20).
- API model ID changes: gpt-6-sol → gpt-6.1-sol (both remain billable; 6 Sol no longer appears in the catalog front page but is still on the pricing page at 2026-10-09).
- NOT officially documented: response-format quirks, tool-calling behavioral parity, stop-token behavior, latency distributions. Treat all of that as UNKNOWN until you test.
Official pricing, every tier
USD per 1M tokens from vendor pricing pages. 'Not published' means exactly that — we don't extrapolate.
| Service tier | GPT-6.1 Sol | GPT-6 Sol |
|---|---|---|
| Standarddefault on-demand serving | in $2 · cached $0.10 · out $10 cache write $2.5 · >272K in/req | in $2 · cached $0.20 · out $10 cache write $2.5 · >272K in/req |
| Batchasync jobs, 50% off | in $1 · cached $0.05 · out $5 cache write $1.25 · >272K in/req | in $1 · cached $0.10 · out $5 cache write $1.25 · >272K in/req |
| Flexslower/cheaper serving, 50% off | in $1 · cached $0.05 · out $5 cache write $1.25 · >272K in/req | in $1 · cached $0.10 · out $5 cache write $1.25 · >272K in/req |
| Fastpriority speed, 2x (ex-Priority) | in $4 · cached $0.20 · out $20 cache write $5 · >272K in/req | in $4 · cached $0.40 · out $20 cache write $5 · >272K in/req |
| Ultrafast6x standard price | in $12 · cached $0.60 · out $60 cache write $15 · >272K in/req | Not published |
| Long-context band | GPT-6.1 Sol | GPT-6 Sol |
|---|---|---|
| Long-context ratesapplies per request above the model threshold | in $4 · cached $0.20 · out $15 cache write $5 · >272K in/req | in $4 · cached $0.40 · out $15 cache write $5 · >272K in/req |
USD per 1M tokens. Every input token is billed in exactly one bucket — uncached, cached read, or cache write; a write replaces the input rate for those tokens (official rule), so write rates are per tier as listed.OpenAI long context = input > 272,000 tokens per single request. · OpenAI pricing (verified 2026-10-09)
Tier ≠ quality change
Run this comparison on your own numbers
Preloaded with a workload typical for this comparison. Change anything — results recompute locally.
| Model | Per request | Monthly |
|---|---|---|
| GPT-6.1 Sol cheapest | $0.0674 | $1,348.00 |
| GPT-6 Sol | $0.0698 | $1,396.00 |
Monthly spread between the cheapest and priciest selected model: $48.00.
Breakdown: GPT-6.1 Sol · standard tier · standard context
- Fresh input (uncached): 0 tok × $2/M$0.00
- Cached input (reads): 24,000 tok × $0.1/M$0.0024
- Cache writes: 6,000 tok × $2.5/M$0.015
- Output (incl. reasoning tokens): 5,000 tok × $10/M$0.05
- Per request$0.0674
- 6,000 tokens billed at the cache-write rate instead of the input rate (mutually exclusive buckets — no double charge).
Rates: official pricing · verified 2026-10-09
- Per-token arithmetic only: it does not model quality differences or cross-vendor output-token-count differences for the same task.
- Reasoning tokens are billed as output tokens at the chosen model's output rate.
- Long-context band is decided per request by total input tokens vs the provider threshold, not by monthly volume.
- Cache accounting uses the official mutually-exclusive buckets: each input token is billed once, as uncached input, cached read, OR cache write — writes replace the input rate for those tokens, never add to it.
- Claude rows bill the 5-minute-TTL cache write; the published 1-hour TTL ($8/MTok standard) is not modeled (no TTL input).
- Rows flagged 'derived' are computed from a stated official rule (e.g. Anthropic batch -50%), not read verbatim from a price table.
- Missing official prices render as Unsupported — we never apply a guessed multiplier (e.g., GPT-6 Sol has no Ultrafast tier).
What the evidence actually says
Independent and vendor-run evals, each with publisher, runner, and harness. Numbers from different runs are never merged.
| Benchmark | Score | Effort | Harness | Evidence |
|---|---|---|---|---|
| DeepSWEv1.1⚠ run comparability unknown | ||||
| gpt-6.1-sol | 75.2 % | high | OpenAI announcement evals | Vendor-reported |
| gpt-6-sol | 68.8 % | best-reported | OpenAI announcement evals (quoted as prior best) | Vendor-reported |
| AutomationBench1.0.6 | ||||
| gpt-6.1-sol | +2.2 pp vs Claude Opus 5.5; +4.8 pp vs GPT-6 Sol | medium | OpenAI announcement (underlying runner not stated — AutomationBench is run/reported by Zapier per Anthropic's footnote) | Vendor-reported |
| OSWorld2.0 offline set, v2026.08.08 | ||||
| gpt-6.1-sol | within 2.1 pp of GPT-6 Astra; +7 pp vs GPT-6 Sol | max | OpenAI announcement evals | Vendor-reported |
| Terminal-Bench-Science cost per task0.1 | ||||
| gpt-6.1-sol | 5.47 usd-per-task | unspecified | OpenAI announcement evals | Vendor-reported |
| Factuality error rateunspecified⚠ run comparability unknown | ||||
| gpt-6.1-sol | 7.7 error-rate-% | low | OpenAI announcement evals | Vendor-reported |
| gpt-6-sol | 11.4 error-rate-% | low | OpenAI announcement evals (quoted baseline) | Vendor-reported |
| GDP.pdf (professional PDF Q&A)unspecified | ||||
| gpt-6.1-sol | scores higher than Opus 5.5 with fallbacks at less than half the cost per task; near Astra at ~1/5 cost per task | multiple settings | OpenAI announcement evals | Vendor-reported |
| AA Intelligence Indexv4.3.2 | ||||
| gpt-6.1-sol | 52 index-points | config 'Sol (Max)' per AA | Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models) | Third-party |
| Terminal-Bench4.0 | ||||
| gpt-6.1-sol | 56 % | — | Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models) | Third-party |
| SciCodeunspecified | ||||
| gpt-6.1-sol | 54 % | — | Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models) | Third-party |
| GDPval-AAv2.1 | ||||
| gpt-6.1-sol | 1592 elo | — | Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models) | Third-party |
| AA-Briefcasev1.1 | ||||
| gpt-6.1-sol | 1557 elo | — | Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models) | Third-party |
| Humanity's Last Examunspecified | ||||
| gpt-6.1-sol | 53 % | — | Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models) | Third-party |
| AutomationBench-AAunspecified | ||||
| gpt-6.1-sol | 65 % | — | Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models) | Third-party |
| AA cost per Intelligence-Index taskv4.3.2 | ||||
| gpt-6.1-sol | 0.72 usd-per-task | — | Artificial Analysis Intelligence Index v4.3.2 (independent, same harness for both models) | Third-party |
Same benchmark name, multiple rows — run comparability unknown
DeepSWE: gpt-6.1-sol = 75.2 (OpenAI announcement evals); gpt-6-sol = 68.8 (OpenAI announcement evals (quoted as prior best)). Same-harness comparison vs gpt-6-sol and gpt-6-astra only.
Factuality error rate: gpt-6.1-sol = 7.7 (OpenAI announcement evals); gpt-6-sol = 11.4 (OpenAI announcement evals (quoted baseline)). Whether these rows share a run is not documented in the sources; treat them as not comparable.
We never average or merge these. See methodology.Sources: OpenAI — Introducing GPT-6.1 Sol (DevDay 2026) · Artificial Analysis (independent) — Independent same-harness comparison: GPT-6.1 Sol vs Claude Opus 5.5 (Intelligence Index v4.3.2)
Which model for which workload
Confidence labels: Measured = same-harness evidence · Inferred = reasoned from documented evidence · Insufficient evidence = no comparable public data.
| Workload | Our pick | Confidence | Why |
|---|---|---|---|
| Daily coding tasks | GPT-6.1 Sol | Measured | Better published evals at identical list price; upgrade is close to free on quality-per-dollar. Run your own regression first; behavior parity is not officially documented. |
| Budget-sensitive / high volume | GPT-6.1 Sol | Measured | Same price, fewer errors at low effort → fewer retries; cache-heavy loops save directly via $0.10 cached reads. |
| Latency-sensitive production | GPT-6.1 Sol | Price-derived | Only 6.1 Sol offers Ultrafast in this pair; Fast tier pricing is identical ($4/$20). |
| Complex repository maintenance | GPT-6.1 Sol | Measured | OSWorld +7 pp at max effort and DeepSWE +6.4 pp at lower effort point the same direction. |
What officially changed (and what didn't)
Only confirmed rows come from vendor pages; 'Unknown' rows are gaps we refuse to fill with assumptions.
| Aspect | Before (gpt-6-sol) | After (gpt-6.1-sol) | Status |
|---|---|---|---|
| Model ID | gpt-6-sol | gpt-6.1-sol | Confirmed |
| Standard input / output price (short ctx) | $2.00 / $10.00 | $2.00 / $10.00 | Confirmed |
| Cached input price (short ctx) | $0.20 | $0.10 | Confirmed |
| Ultrafast tier | not offered | $12 / $60 (short ctx) | Confirmed |
| Reasoning efforts | low / medium / high / xhigh / max | low / medium / high / xhigh / max | Confirmed |
| Prompt/tool behavior parity | — | not officially documented | Unknown |
| Deprecation of gpt-6-sol | — | no timeline announced as of 2026-10-09 | Unknown |
When staying on the old model is reasonable
- You need frozen, reproducible outputs for compliance or benchmarking (no vendor guarantees behavioral freezes — freezing means 'don't change anything', including the model ID).
- Your regression suite hasn't run yet — that's a reason to wait days, not months: the upgrade is same-price on standard tier.
- You depend on a third-party gateway that hasn't enabled gpt-6.1-sol (check with the provider; don't assume).
GPT-6 Sol → GPT-6.1 Sol migration checklist
Behavioral compatibility is not officially documented, so the checklist is: freeze a golden set, swap the ID in staging, measure, then roll out gradually. Progress is saved in your browser only and can be exported.
GPT-6 Sol → GPT-6.1 Sol migration checklist
0/8 completeBehavioral compatibility is not officially documented, so the checklist is: freeze a golden set, swap the ID in staging, measure, then roll out gradually. Progress is saved in your browser only and can be exported.
State is stored in your browser's localStorage only — nothing is uploaded, and the export file is generated locally.
Frequently asked questions
Is gpt-6-sol being deprecated?
No timeline has been announced as of 2026-10-09. It was removed from the model catalog front page but remains on the official pricing page and fully billable. We monitor the OpenAI changelog and update this page when that changes.
Will my prompts behave identically after switching the model ID?
That is not officially documented. Published evals show quality improvements, which by definition means behavior changed. Freeze a golden set of your real tasks, run both IDs, and diff the outputs — the checklist above walks you through it.
Does upgrading change my bill?
Standard list prices are identical ($2/$10). Two real effects: cached input halves to $0.10/MTok (cache-heavy agent loops get cheaper), and the optional Ultrafast tier ($12/$60) adds a new way to spend — only if you opt in. If 6.1 Sol reaches your quality bar at a lower reasoning effort, per-task cost can drop further.
What did OpenAI officially confirm as different?
From the GPT-6.1 Sol announcement: better DeepSWE v1.1 (+6.4 pp over 6 Sol's best, at lower effort), factuality error rate 11.4% → 7.7% at low effort, OSWorld +7 pp at max, cached input 50% cheaper, Ultrafast tier availability. Everything else — tool-calling parity, latency, formatting — is unconfirmed.
Related
Sources & freshness
- OpenAI — OpenAI API model catalog (model IDs, context windows, reasoning efforts, capabilities)https://developers.openai.com/api/docs/models · accessed 2026-10-09GPT-6 Sol is NOT listed in the catalog front page at access time (still priced on the pricing page). Knowledge cutoffs: Apr 30 2026 (Astra/6.1 Sol), May 18 2026 (Luna).
- OpenAI — OpenAI API pricing (GPT-6 family tier matrix: standard/batch/flex/fast/ultrafast, short/long context, cache)https://developers.openai.com/api/docs/pricing · accessed 2026-10-09Long context defined as >272K input tokens per request (2x input and cache rates, 1.5x output for the whole request). 'Priority processing' renamed 'Fast mode' on 2026-07-30. Flex and Batch are half of Standard for the GPT-6 family, Fast is 2x, Ultrafast is 6x. Cache writes cost 1.25x the uncached input rate, reads 0.1x (0.05x on GPT-6.1 Sol). Cache-write rates are captured on every published tier row, not Standard only.
- OpenAI — Introducing GPT-6.1 Sol (DevDay 2026)https://openai.com/index/introducing-gpt-6-1-sol · published 2026-09-29 · accessed 2026-10-09Self-reported evals (OpenAI harness). Includes per-task cost quotes (Terminal-Bench-Science). Announces GPT-6.1 Sol Ultrafast 'coming soon' in Codex. Availability: ChatGPT Work + Codex, not yet in Chat.
- Artificial Analysis (independent) — Independent same-harness comparison: GPT-6.1 Sol vs Claude Opus 5.5 (Intelligence Index v4.3.2)https://artificialanalysis.ai/models/comparisons/gpt-6-1-sol-vs-claude-opus-5-5 · accessed 2026-10-10Independent same-harness evaluation of both models. Cost-per-task = weighted average per Intelligence Index task; reflects measured token verbosity. Effort configs as labeled by AA ('Sol (Max)' vs 'Opus 5.5 (Max, Default Fallback)').
Pricing data last verified 2026-10-09. OpenAI ships roughly weekly — re-verify before making spend commitments. See the methodology page for update cadence and stale-data flags.