OpenAI replaced GPT-5.6 Sol with GPT-6 Sol on 22 September 2026 at exactly half the price. Artificial Analysis ran both models through all six effort levels in the same harness, and the result is the cleanest generational comparison of the year: every rung of the ladder scored within a point of the model it replaced, and every rung cost roughly half as much to run.

TL;DR

This is a price cut wearing a model launch. Across six effort levels on Artificial Analysis Intelligence Index v4.3.2, GPT-6 Sol scores between 0.1 and 0.9 points above GPT-5.6 Sol at the same setting — an average of half a point — while cost per task drops 47–55% at every rung. The independent summary is blunt: "Intelligence Index and Coding Agent Index scores remain level with GPT-5.6, with progress in some evaluations and regressions in others."

  • The ladder didn't move; the price did. Medium: 39 → 39.8 for $0.50 → $0.25. Max: 47 → 47.5 for $1.99 → $1.06.
  • Coding is the one real capability gain. Coding Agent Index 55 → 57, Terminal-Bench 4.0 37% → 43%, SWE-Atlas-QnA 54% → 58% — at about half the cost per coding task.
  • Knowledge work went backwards. GDPval-AA v2.1 loses about 100 Elo. Artificial Analysis traces it to "shorter deliverables that more often omit required elements."
  • Hallucinations fell from 92% to 60% — but partly by declining to answer. Questions attempted dropped from 99% to 83%, and accuracy fell 5 points to 54%.
  • The cliff: GPT-5.6 Sol's $4/$20 rate is promotional only, guaranteed "at least through November 21, 2026." Its list price is $5/$30. GPT-6 Sol's $2/$10 is permanent.
+0.5
average index gain
across six effort levels
−51%
average cost per task
at matched settings
55 → 57
Coding Agent Index
the one real gain
−100
GDPval-AA Elo
the one real loss

Two generations, five months of repricing

The GPT-5.6 Sol you might be running today is not priced the way it launched. It is worth walking the timeline before comparing anything, because two of the three price movements in this story belong to the old model.

DateWhat changedEffect on Sol
9 July 2026GPT-5.6 Sol, Terra and Luna reach general availability after a preview that opened 26 JuneLaunches at $5 / $30 per million tokens — GPT-5.5's exact rate
30 July 2026OpenAI passes its own efficiency gains to the cheaper tiersTerra and Luna are cut. Sol is left alone.
21 August 2026Sol's turn: 20% off input, 33% off output$4 / $20 — described by OpenAI as promotional, available "at least through November 21, 2026"
3 September 2026GPT-6 Astra ships at $10 / $50 and takes the flagship slotSol becomes the working tier rather than the frontier one
22 September 2026GPT-6 Sol and Luna replace both GPT-5.6 tiersSol at $2 / $10. An OpenAI spokesperson confirmed to VentureBeat these are permanent rates, "not promotional or introductory pricing"

That last row is the whole story. Against the promotional rate, GPT-6 Sol is exactly 50% cheaper in both directions. Against GPT-5.6 Sol's list rate of $5 / $30 it is 60% cheaper on input and 67% cheaper on output. And unlike the discount it replaces, the new rate has no review date attached.

Spec sheet

SpecificationGPT-6 SolGPT-5.6 SolChange
Released22 September 20269 July 2026+75 days
API model IDgpt-6-solgpt-5.6-sol the gpt-5.6 alias routes here
Input$2.00 / MTok$4.00 promo · $5.00 list−50% promo, −60% list
Cached input$0.20 / MTok$0.40 promo−50%
Cache write$2.50 / MTok$5.00 promo−50%
Output$10.00 / MTok$20.00 promo · $30.00 list−50% promo, −67% list
Over 272K input, in / out$4 / $15$8 / $30−50%
Batch and Flex50% off — $1 / $5$2 / $10−50%
Fast mode2× — $4 / $20$8 / $40−50%
Context window1.05M tokens1.05M tokensunchanged
Max output128K tokens128K tokensunchanged
Knowledge cutoff20 April 202616 February 2026+2 months
Reasoning controlnone, low, medium, high, xhigh, max — default mediumnon-reasoning, low, medium, high, xhigh, maxrenamed, same shape
Modalitiestext + images → texttext + images → textunchanged
Where it runsOpenAI API, Codex, ChatGPT Work+ Amazon Bedrock, Microsoft Azurenarrower on day one
System cardnone published at launchpublished with the GPT-5.6 familyregression
Promotional-cliff risknone — permanent rate$4/$20 through 21 Nov 2026 at leastGPT-6 Sol

The pricing card is the same card, halved. Artificial Analysis notes both generations carry "the same 90% discount for cache reads and 25% premium for cache writes." Cache read is exactly 10% of the input rate and cache write exactly 125% of it, on both models. That means the price cut is linear: there is no line item on the invoice where the old model is competitive.

The ladder, rung by rung

This is the comparison that matters, because it is the only one where both models ran the identical ten evaluations under one operator at every effort level. Artificial Analysis Intelligence Index v4.3.2 bundles AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR v1.1. Costs are per task at list prices, including cache reads and writes.

EffortGPT-5.6 SolGPT-6 SolScore deltaCost delta
noneGPT-5.6 Sol's equivalent is "non-reasoning"28 · —28.1 · $0.33+0.1n/a
low33 · $0.2633.9 · $0.13+0.9−50%
mediumboth models' default39 · $0.5039.8 · $0.25+0.8−50%
high42 · $0.8142.8 · $0.37+0.8−54%
xhigh44 · $1.1844.1 · $0.53+0.1−55%
max47 · $1.9947.5 · $1.06+0.5−47%

Read the two columns as a step function that didn't step. Six rungs, six moves of less than a point, and an average cost reduction of 51%. The family ceiling moves from 47 to 48 — a single point across a seventy-five-day release cycle that also included a flagship launch in between.

GPT-6 Sol · low 33.9 $0.13 GPT-6 Sol · medium 39.8 $0.25 GPT-5.6 Sol · low 33 $0.26 GPT-6 Sol · high 42.8 $0.37 GPT-5.6 Sol · medium 39 $0.50 GPT-6 Sol · xhigh 44.1 $0.53 GPT-5.6 Sol · high 42 $0.81 GPT-6 Sol · max 47.5 $1.06 GPT-5.6 Sol · xhigh 44 $1.18 GPT-5.6 Sol · max 47 $1.99 GPT-6 Sol GPT-5.6 Sol Sorted by cost per Intelligence Index task · bar length = index score
Every rung of the new ladder slots between two rungs of the old one. GPT-6 Sol at xhigh scores 44.1 for $0.53; GPT-5.6 Sol needs $1.18 to reach 44. GPT-6 Sol's most expensive setting, max at $1.06, costs less than GPT-5.6 Sol's second-most-expensive setting at $1.18 — and scores 3.5 points higher. Source: Artificial Analysis Intelligence Index v4.3.2.

Two rungs deserve a second look. GPT-6 Sol at low costs $0.13 — half of GPT-5.6 Sol's cheapest setting — and already outscores it by 0.9 points. And GPT-6 Sol at medium, $0.25, undercuts GPT-5.6 Sol at low while scoring 6.8 points higher. If you are on the old model and only ever wanted a cheap setting, the new model's third-cheapest rung beats your cheapest rung for less money.

Where it improved: coding

The one place GPT-6 Sol posts a gain you can feel is agentic software engineering. In OpenAI's Codex harness — the same harness for both models — Artificial Analysis's Coding Agent Index puts GPT-6 Sol at 57 against 55, and the components show where it came from.

EvaluationGPT-6 Sol (max)GPT-5.6 Sol (max)Delta
Artificial Analysis Coding Agent IndexCodex harness5755+2
Terminal-Bench 4.0agentic terminal work43%37%+6
SWE-Atlas-QnAcodebase question answering58%54%+4
Cost per coding-agent task$2.99~$6.00−50%

Six points on Terminal-Bench 4.0 is the largest single improvement across the whole release, and it lands on the benchmark that best predicts whether an agent can finish a job in a shell without you watching. The cost line matters just as much: the index gain is real, and you pay half as much for it.

Worth keeping in perspective on both sides. On the public Terminal-Bench 4.0 leaderboard, GPT-5.6 Sol sits at 37.3% and GPT-6 Astra at 57.7% — so GPT-6 Sol has pulled the working tier to within striking distance of where the frontier was two weeks earlier, without touching the frontier itself.

Where it regressed: knowledge work

Artificial Analysis found two knowledge-work regressions, and they are the reason the Intelligence Index didn't move despite the coding gains:

  • GDPval-AA v2.1 falls about 100 Elo. This is the benchmark adapted from OpenAI's own dataset of economically valuable tasks across 44 occupations, so a regression here is not a harness artefact.
  • AA-Briefcase v1.1 is level — no gain, no loss, on a private evaluation across multi-week knowledge-work projects with thousands of input files.

Their diagnosis is specific and worth quoting, because it describes a failure mode you can check for in your own evals: the regressions "tend to be driven by reduced presentation quality and deliverables that omit rubric elements." Artificial Analysis notes its team "has manually inspected hundreds of model outputs" to reach that conclusion.

Translated: the new model produces shorter deliverables that skip required components. On an agentic coding task that is invisible. On a 30-page analysis where the client expects six specific sections, it is the whole job.

Hallucinations half, accuracy down

Both generations hallucinate a lot on AA-Omniscience — a benchmark that rewards correct answers, penalises wrong ones, and does not penalise refusing to answer. GPT-6 Sol cuts the hallucination rate from 92% to 60%. That is a large improvement, and it comes with a catch that Artificial Analysis states plainly.

GPT-6 Sol · max GPT-5.6 Sol · max Hallucination rate — lower is better 60% 92% Accuracy on questions attempted 54% 59% Share of questions attempted 83% 99%
The hallucination cut is partly an abstention cut. GPT-6 Sol answers 16 percentage points fewer questions, gets 5 points fewer of those right, and produces a wrong answer 32 points less often. On the composite AA-Omniscience Index — which scores knowledge reliability rather than raw accuracy — it improves from 22 to 27. Source: Artificial Analysis.

Is that a good trade? For most production use, yes: a model that says "I don't know" costs you a retry, while a model that invents a figure costs you a wrong decision. But it is not free. If your pipeline depends on the model attempting everything and you have downstream validation, you have traded away 5 points of coverage-adjusted accuracy for a safer failure mode. Measure it on your own traffic rather than assuming.

What OpenAI's charts say, and what they don't

OpenAI's launch page benchmarks GPT-6 Sol against Claude Opus 5, not against GPT-5.6 Sol on most of its charts, and it uses Zapier's and Cognition's runs rather than Artificial Analysis's. Two of its numbers are worth separating from the independent ones.

OpenAI's claimFigureRead it as
AutomationBenchGPT-6 Sol (xhigh) 33.2% at $0.27 per task; Claude Opus 5 (max) 26.9% at 11.1× the costA real win, but a different benchmark version from Artificial Analysis's AutomationBench-AA, where the same pair reads 62% vs 60%
Agents' Last ExamGPT-6 Sol (max) 56.4%, "above Opus 5's highest score at 60% lower cost per task"Effort level of the Opus 5 comparison is not stated, so the cost claim is not reproducible
DeepSWE v1.1GPT-6 Sol (max) 68.8% vs Claude Fable 5 (xhigh) 69.9%, at ~80% lower cost0.6 points behind the leader at the effort setting OpenAI chose to quote
OSWorld 2.0 offlineGPT-6 Sol (xhigh) 60.5% vs Opus 5 (medium) 60.3%OpenAI's own summary: the new model at xhigh is "similar" to the older Claude model at medium
Internal factuality eval"About half as many mistakes as its predecessor"Built from de-identified ChatGPT conversations where users flagged factual errors — consistent with the AA-Omniscience direction
Restriction circumventionFalls from 68.2% to 64.4% on adversarial runs containing explicit "access denied" signalsA small improvement. OpenAI stresses these are deliberately difficult, low-stakes tests without product safeguards

The pattern is consistent: GPT-6 Sol is a modest capability bump at half the price, and OpenAI's strongest charts measure cost per task rather than raw score, because that is where the release actually moves.

What half price is worth on a real invoice

Because every line item on the card is scaled by exactly 0.5, the arithmetic is unusually clean. Here is the illustrative cache-heavy agent task — 8M cache reads, 400K uncached input, 600K cache writes, 300K output — priced at all three rate cards. Token counts are held identical; only the published rates change.

LineGPT-6 SolGPT-5.6 Sol (promo)GPT-5.6 Sol (list)
Cache reads · 8,000,000$1.60$3.20$4.00
Uncached input · 400,000$0.80$1.60$2.00
Cache writes · 600,000$1.50$3.00$3.75
Output · 300,000$3.00$6.00$9.00
Total$6.90$13.80$18.75
Versus GPT-6 Sol2.0× more2.7× more

Real token counts differ between the models, and Artificial Analysis's measured figures already account for that — GPT-6 Sol actually uses more output tokens per task, 31k against 29k. The measured result is still a 47% cost reduction at max effort, because the price cut is larger than the token increase.

The upgrade path, and the November cliff

Nothing forces your hand today, and that is worth stating precisely because the promotional pricing does not last forever.

  • Neither model has a shutdown date. OpenAI's deprecations page lists no retirement for gpt-5.6-sol. Microsoft's Foundry schedule shows both gpt-5.6-sol and gpt-5.6-luna retiring on 11 January 2028.
  • GPT-5.6 Sol is still the migration target for older models. o1, o3, o3-pro and o3-deep-research are deprecated and point at gpt-5.6-sol — not GPT-6 Sol. The original GPT-5 and GPT-5-mini retire on 11 December 2026 pointing at gpt-5.6-sol and gpt-5.6-terra.
  • The promotional rate has a review date. OpenAI guarantees $4/$20 "at least through November 21, 2026" and has said nothing about afterwards. If it reverts to list, this task goes from $13.80 to $18.75 — a 36% increase on a model that didn't change.

The clean way to think about the cliff: GPT-6 Sol costs less at every effort level than GPT-5.6 Sol's promotional rate — its two cheapest rungs undercut the old model's cheapest, and its most expensive rung at $1.06 beats the old model's $1.18. There is no setting where sticking with GPT-5.6 Sol saves money today, and there is a dated risk where it costs 36% more in December.

Verdict

The short version

Migrate, and expect the migration to be a budget line rather than a capability leap. GPT-6 Sol is not a better model in any way you will notice on a benchmark table — half a point of index, two points of coding index, and a knowledge-work regression that will show up in long deliverables. It is a better purchase: the same quality band for half the price, with no expiry date.

The exception is coding. Terminal-Bench 4.0 moving six points and the Coding Agent Index moving two is a genuine improvement on the benchmark that matters most for unattended agents, and it arrives at half the cost per coding task. If you run Codex or a terminal agent, this is a straight upgrade. If you run long-form knowledge work, test the deliverable completeness before you switch — that is exactly where the regression lives.

Your workloadRoute toWhy
Terminal agents, Codex, CLI automationGPT-6 SolTerminal-Bench 4.0 +6 points, Coding Agent Index +2, at half the cost per task
High-volume, cost-constrained productionGPT-6 SolEvery rung undercuts the old model's equivalent; the rate is permanent
Workloads pinned to promotional pricingGPT-6 SolRemoves the 21 November cliff from your model
Long-form deliverables with strict rubricsTest firstGDPval-AA lost ~100 Elo on omitted rubric elements and weaker presentation
Pipelines needing Bedrock or Azure on day oneStay on GPT-5.6 SolGPT-6 Sol launched API, Codex and ChatGPT Work only
Workloads needing reasoning fully offEitherBoth offer it; GPT-6 Sol's none costs $0.33 and scores 28.1
Auditing a model against a published system cardGPT-5.6 SolOpenAI published no system card for GPT-6 Sol or Luna at launch
Absolute frontier capabilityNeither — GPT-6 Astra53 on the same index, at $10/$50

FAQ

Is GPT-6 Sol actually smarter than GPT-5.6 Sol?

Marginally, and only on some evaluations. At matched effort settings it gains 0.1–0.9 index points, averaging half a point, and the family ceiling moves from 47 to 48. Coding improved meaningfully (Terminal-Bench 4.0 +6, SWE-Atlas-QnA +4); knowledge work regressed badly (GDPval-AA −100 Elo). Artificial Analysis's own summary is "level with GPT-5.6, with progress in some evaluations and regressions in others."

Why is GPT-6 Sol exactly half the price?

OpenAI says "improvements to inference and caching" let it cut prices while increasing capability, and an OpenAI spokesperson confirmed to VentureBeat that the rates are permanent rather than promotional. The halving is uniform: input, output, cache reads, cache writes, long-context rates, batch and fast mode are all scaled by 0.5 against the promotional card.

What happens on 21 November 2026?

Nobody knows yet. OpenAI's documentation guarantees GPT-5.6 Sol's $4/$20 promotional rate "at least through November 21, 2026" and has not said what follows. Against the list rate of $5/$30, the illustrative agent task in this article goes from $13.80 to $18.75 — a 36% increase on an unchanged model. GPT-6 Sol's $2/$10 carries no such date.

Does the lower hallucination rate mean it is more accurate?

Not exactly. GPT-6 Sol answers fewer questions — 83% against 99% — and gets 54% of those right against 59%. What improved is reliability in the sense that matters for production: it produces a wrong answer far less often. On AA-Omniscience's composite index, which rewards correct answers and penalises hallucinations with no penalty for abstaining, GPT-6 Sol improves from 22 to 27.

Should I switch off max effort to save money?

Probably yes, unless you can point at a specific failure that only max fixes. On the old model the last rung from xhigh to max cost 69% more for 3 points. On GPT-6 Sol, xhigh gives 44.1 for $0.53 and max gives 47.5 for $1.06 — double the cost for just over three points. high at $0.37 is the best value rung if your work tolerates 42.8.

Run both models side by side on your own tasks

Benchmarks decide direction; your own eval decides the switch. Put three tasks from last sprint through GPT-6 Sol at high and max and compare the receipts — that is the number the ladder can't tell you.

Open a new chat on CodingFleet →

Sources & further reading

Benchmark scores are vendor- or evaluator-reported and labelled as such. Cross-vendor numbers come from different harnesses, effort settings and tool configurations and are directional. Where two published figures disagree — for example AutomationBench run by Zapier versus AutomationBench-AA run by Artificial Analysis — the article states both rather than picking one. Prices are USD per million tokens as of late September 2026.