OpenAI shipped GPT-6 Sol and GPT-6 Luna on 22 September 2026, roughly ninety minutes after Anthropic published Claude Opus 5.5. Sol lists at $2 / $10 per million tokens — half of Opus 5.5's $4 / $20. Neither vendor compared the two directly: OpenAI's launch charts show Opus 5, and Anthropic's table shows GPT-5.6 Sol. So the comparison below leans on the one harness that ran both models on the same tests at every effort level.
TL;DR
GPT-6 Sol is the cheapest way to reach a good score, and Claude Opus 5.5 is the only way to reach a great one. On Artificial Analysis's Intelligence Index v4.3.2, Sol is cheaper than any Opus 5.5 setting for every score up to about 44. Above that it runs out of headroom: its best result, 47.5 at max for $1.06 a task, sits below Opus 5.5's default medium setting, 51.2 for $1.34.
- Price: $2 / $10 versus $4 / $20 per million tokens — but cache reads cost the same $0.20 on both, and only Sol charges a long-context surcharge.
- Coding is the widest gap: Opus 5.5 at
mediumscores 52.5% on Terminal-Bench 4.0 against Sol's 43.9% atmax, and reaches 59.6% atxhigh. Sol falls to 26.3% athigh. - Knowledge work favours Opus 5.5: +89 Elo on GDPval-AA and +159 Elo on AA-Briefcase at matched settings.
- Two rows are level: AutomationBench-AA (61.2% vs 61.6%) and long-context reasoning (84.3% vs 83.7%). On business workflows Sol matches Opus 5.5 at
mediumfor about 40% of the cost. - The catch: Sol's
maxsetting makes the two models cost almost the same, and on two of OpenAI's own charts the older GPT-5.6 Sol still scores higher.
$2/$10 vs $4/$20
Sol has no reach
vs Sol's best, 47.5
identical on both
Two launches, one day, opposite strategies
OpenAI positioned Sol and Luna as Astra's training methods brought down-market, with the savings handed to users. GPT-6 Astra, released 3 September at $10 / $50, stays the flagship for "the most demanding and important projects." Sol replaces GPT-5.6 Sol at half the price; Luna replaces GPT-5.6 Luna at half the price. Terra, the middle tier, is not part of this generation — OpenAI's GPT-6 pricing section lists only Astra, Sol and Luna.
Both survivors take over the names of the models they replace, which makes version numbers load-bearing: gpt-5.6-sol is still a live model on promotional pricing through at least 21 November 2026, and gpt-6-sol is a different thing entirely. Developers first saw gpt-6-sol where they expected gpt-5.6-sol in live API traffic on 16 September, six days before the launch.
Anthropic's approach on the same day was the opposite: keep the price band, move the capability. Opus 5.5 cut tokens 20%, cache reads 60% and cost per task 40% against Opus 5 — and did it while moving up on every published benchmark. OpenAI's release moves the price down and, on two of its own coding charts, the score with it.
Spec sheet
| Specification | Claude Opus 5.5 | GPT-6 Sol | GPT-6 Luna |
|---|---|---|---|
| Released | 22 September 2026 | 22 September 2026 ~90 min later | 22 September 2026 |
| API model ID | claude-opus-5-5 | gpt-6-sol | gpt-6-luna |
| Input | $4 / MTok | $2 / MTok | $0.10 / MTok |
| Cached input | $0.20 / MTok | $0.20 / MTok | $0.01 / MTok |
| Cache write | $5.00 (5 min) · $8.00 (1 h) | $2.50 | $0.125 |
| Output | $20 / MTok | $10 / MTok | $0.50 / MTok |
| Above 272K input | no surcharge | 2× input & cache, 1.5× output | 2× input & cache, 1.5× output |
| Batch / Flex | $2 / $10 | $1 / $5 | $0.05 / $0.25 |
| Fast mode | $8 / $40 | $4 / $20 | $0.20 / $1 |
| Context window | 1M tokens | 1.05M tokens | 1.05M tokens |
| Max output | 128K (300K batch beta) | 128K | 128K |
| Reasoning off | not offered — low is the floor | effort: none | effort: none |
| Effort levels | low → max | none, low, medium, high, xhigh, max | none, low, medium, high, xhigh, max |
| Default effort | medium | medium | medium |
| Knowledge cutoff | June 2026 | April 20, 2026 | May 18, 2026 |
| Modalities | text + images → text | text + images → text | text + images → text |
| Where it runs | Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud, Microsoft Foundry | OpenAI API only | OpenAI API only |
| System card | published | none at launch | none at launch |
Sol's rates are Claude Sonnet 5's rates. Input $2, short cache write $2.50, cache read $0.20, output $10 — line for line identical to Sonnet 5 on Anthropic's pricing page. Luna at $0.10 / $0.50 is a tenth of Claude Haiku 4.5's $1 / $5, which puts it in territory where hosted open-weight models usually live.
The ladder: the same tests, at every effort level
This is the most useful table in the comparison, because it is the only one where both models ran the identical ten evaluations under one operator. Artificial Analysis's Intelligence Index v4.3.2 averages agentic coding, business workflows, knowledge-work documents, science, long-context reasoning and factual knowledge. Costs are per task at list prices, including cache reads and writes.
| Effort | Claude Opus 5.5 | GPT-6 Sol |
|---|---|---|
| nonereasoning switched off | not offered | 28.1 · $0.33 |
| low | 42.3 · $0.55 | 33.9 · $0.13 |
| mediumboth models' default | 51.2 · $1.34 | 39.8 · $0.25 |
| high | 53.6 · $1.82 | 42.8 · $0.37 |
| xhigh | 56.0 · $3.46 | 44.1 · $0.53 |
| max | 57.6 · $5.98 | 47.5 · $1.06 |
max costs less than the cheapest Opus 5.5 setting, and Sol at xhigh — 44.1 for $0.53 — already beats Opus 5.5 at low, 42.3 for $0.55. Everything Opus 5.5 does from medium upward scores higher than anything Sol can reach. The useful question is not which model is better but which price band the task belongs in.One result goes against intuition. Sol with reasoning switched off scored 28.1 for $0.33 a task — below Sol at low (33.9) and more than twice as expensive. Almost all the extra cost is input, and the likely reason is that a model without reasoning takes more steps to finish, and every step re-sends the conversation. Turning reasoning off is not a reliable way to save money on agent work.
Test by test: where the gap opens and where it closes
The fairest pairing is each model at the setting a team would actually use for demanding work: Opus 5.5 at its default medium against Sol at max, which costs about the same. Sol at xhigh is included as the cost-conscious option.
| Evaluation | Opus 5.5 · medium | Sol · max | Sol · xhigh |
|---|---|---|---|
| Cost per index task | $1.34 | $1.06 | $0.53 |
| Output tokens per task | 25.7k | 31.2k | 16.0k |
| Intelligence Index v4.3.2 | 51.2 | 47.5 | 44.1 |
| Terminal-Bench 4.0agentic coding | 52.5% | 43.9% | 30.3% |
| AutomationBench-AAbusiness workflows · partial credit | 61.2% | 61.6% | 61.7% |
| GDPval-AAknowledge work · Elo | 1,576 | 1,487 | 1,437 |
| AA-Briefcasework documents · Elo | 1,642 | 1,483 | 1,364 |
| Humanity's Last Exam | 54.7% | 47.9% | 46.3% |
| SciCodescientific coding | 59.3% | 57.6% | 55.1% |
| CritPtphysics research | 27.7% | 30.9% | 28.0% |
| AA-LCRlong-context reasoning | 84.3% | 83.7% | 81.3% |
| AA-Omniscience accuracy | 64.5% | 54.5% | 53.8% |
| AA-Omniscience hallucination ratelower is better | 68.4% | 60.1% | 58.9% |
Coding is the widest gap
On Terminal-Bench 4.0, Opus 5.5 at medium scores 52.5% against Sol's 43.9% at max — and Opus 5.5 climbs to 59.6% at xhigh. Sol falls away sharply at lower settings: 30.3% at xhigh and 26.3% at high. For coding agents that run in a terminal, Sol's lower price buys noticeably weaker output, and the setting you pick matters more than the model you pick.
Worth keeping in perspective on both sides: on the public Terminal-Bench 4.0 leaderboard, GPT-5.6 Sol sits at 37.3%, Claude Opus 5 at 51.8%, Fable 5.1 at 55.8% and GPT-6 Astra at 57.7%. Harnesses differ, but the ordering across Anthropic's current lineup and OpenAI's older flagship is consistent with the Artificial Analysis run.
Knowledge work favours Opus 5.5, with a caveat
On GDPval-AA and AA-Briefcase — benchmarks that rate real work documents in head-to-head comparisons scored as Elo — Opus 5.5 at medium leads Sol at max by about 89 and 159 Elo points. At max, Opus 5.5 reaches 1,846 on GDPval-AA, the same figure Anthropic publishes in its launch table.
Business workflows and long context are level
On AutomationBench-AA — Artificial Analysis's partial-credit run of business tasks across apps — Sol at xhigh (61.7%) matches Opus 5.5 at medium (61.2%) at about 40% of the average cost per task. Long-context reasoning is close too, 83.7% against 84.3%. Sol also leads on CritPt, a physics research benchmark, 30.9% to 27.7%. If your workload is SaaS automation, document retrieval or physics-style analysis, Sol at xhigh is the value pick — and the gap to Opus 5.5 is inside noise.
Knowledge versus guessing
AA-Omniscience rewards correct answers, penalises wrong ones, and does not penalise declining to answer. Opus 5.5 answers more questions correctly — 64.5% against Sol's 54.5% — but when it does not know, it produces a wrong answer 68.4% of the time, against 60.1% for Sol at max and 50.7% at Sol at low. Sol knows less but guesses less. Neither rate is low enough to skip verification on factual work, and the reversal is a useful reminder that Anthropic's calibration advantage is not automatic.
The invoice: half the token price is not half the bill
| Line, per million tokens | Opus 5.5 | GPT-6 Sol | Sol as % of Opus |
|---|---|---|---|
| Input | $4.00 | $2.00 | 50% |
| Cache read | $0.20 | $0.20 | 100% |
| Cache write (Anthropic: 5-minute) | $5.00 | $2.50 | 50% |
| Output | $20.00 | $10.00 | 50% |
| Request over 272K input, in / out | $4 / $20 | $4 / $15 | 100% / 75% |
| Batch, in / out | $2 / $10 | $1 / $5 | 50% |
| Fast mode, in / out | $8 / $40 | $4 / $20 | 50% |
Two lines break the half-price rule. Cache reads cost the same $0.20 per million tokens on both models, and for an agent, cache reads are usually the largest share of the bill. And Anthropic serves the full 1M-token window at standard rates, while OpenAI bills any request over 272K input tokens at twice the input and cache rates and 1.5× the output rate. A single oversized request puts Sol at $4 / $15.
Here is the same illustrative agent task Anthropic used in its Opus 5.5 launch material — 8M cache reads, 400K uncached input, 600K cache writes, 300K output — priced at both rate cards. The token counts are held identical for the comparison; the rates are published ones.
| Illustrative task | Opus 5.5 | GPT-6 Sol | Sol, long prompts |
|---|---|---|---|
| Cache reads · 8,000,000 | $1.60 | $1.60 | $3.20 |
| Uncached input · 400,000 | $1.60 | $0.80 | $1.60 |
| Cache writes · 600,000 | $3.00 | $1.50 | $3.00 |
| Output incl. reasoning · 300,000 | $6.00 | $3.00 | $4.50 |
| Total | $12.20 | $6.90 | $12.30 |
At their defaults the two measured figures are stark: Sol costs $0.25 a task against Opus 5.5's $1.34, about a fifth — but for an index score of 39.8 against 51.2. The cheaper model is cheap because it is doing less work, and the moment you turn Sol up to reach Opus 5.5's default score, you are paying $1.06 for 47.5 and still falling short.
Why the vendors' shared rows don't line up
Both launch posts published results for benchmarks the two models share. They point the same way as the independent run, but they are not a clean comparison — and in one case they are not even internally consistent.
| Benchmark | Anthropic reports | OpenAI reports | Read it as |
|---|---|---|---|
| AutomationBenchZapier-run business workflows | Opus 5.5 40.0% at max (Zapier early access) | Sol 33.2% at xhigh — the best level for Sol | Both are Zapier's own scores, which run far below Artificial Analysis's partial-credit version of the same benchmark (61.2% vs 61.6%) |
| FrontierCode 1.1 MainCognition | Opus 5.5 54.4%; Opus 5 48.0% | Sol 49.3% at max; Opus 5 53.4% | Inconclusive. The vendors disagree about the same model by 5.4 points — a wider spread than the gap between the two models |
| OSWorld 2.0computer use | Opus 5.5 81.8% (partial credit) | Sol 64.4% (offline set) | Not comparable — different task sets and different scoring |
There is a simpler point underneath: Anthropic's launch table lists GPT-5.6 Sol, not GPT-6 Sol — 37.3% on Terminal-Bench 4.0, 47.5% on FrontierCode, 41.7% on CursorBench, 22.4% on Terminal-Bench-Science 0.1, 28.8% on AutomationBench and 1,588 Elo on GDPval-AA. Those rows describe the model Sol replaced, published the same day OpenAI shipped its replacement. Meanwhile OpenAI's charts benchmark against Opus 5 rather than Opus 5.5, for the same reason. Both vendors had a fresh number available and both chose the older comparison.
OpenAI's own chart data, read carefully
OpenAI's launch page is built on interactive cost-per-task charts at five effort levels. Reading the underlying values rather than the quoted points produces a fairer picture of what Sol and Luna are.
| Benchmark | GPT-6 Sol best | GPT-6 Luna best | Top of chart |
|---|---|---|---|
| AutomationBenchbusiness workflows | 33.2% · xhigh · $0.27 | 20.7% · max · $0.04 | 41.4% · Astra max |
| Agents' Last Examprofessional work | 56.4% · max · $2.93 | 50.9% · max · $0.15 | 59.3% · Astra max |
| FrontierCode 1.1mergeable code | 49.3% · max · $2.14 | 42.4% · max · $0.11 | 53.4% · Opus 5 medium |
| DeepSWE 1.1long software tasks | 68.8% · max · $2.74 | 66.6% · max · $0.22 | 74.1% · Astra xhigh |
| OSWorld 2.0 offlinecomputer use | 64.4% · max · $3.25 | 52.7% · max · $0.27 | 73.5% · Astra max |
| Factual error ratelower is better | 4.5% · xhigh · $0.13 | 7.6% · max · $0.01 | 3.9% · Astra high |
The headline claims check out against OpenAI's own data. On AutomationBench, Sol at xhigh scores 33.2% for $0.27 a task, above Claude Opus 5 at max (26.9%) for a small fraction of its cost. On Agents' Last Exam, Sol at max (56.4%) edges Opus 5's best (55.9% at high) for about 40% of the cost. On OSWorld, Sol at xhigh matches Opus 5 at medium — 60.5% against 60.3% — for roughly a sixth of the cost. And factual reliability improved at every effort level: on OpenAI's internal test of conversations where users had flagged a mistake, Sol's error rate is roughly half GPT-5.6 Sol's at each setting (5.1% against 10.8% at high). OpenAI notes those conversations were selected because they caused errors, so everyday rates are lower.
Four things the charts show that the text does not.
- GPT-5.6 Sol still has the higher top score on two charts. On DeepSWE it scores 72.7% against GPT-6 Sol's 68.8%; on OSWorld, 66.2% against 64.4%. The older model costs more than twice as much per task at those settings, so it only makes sense where the extra points are worth the price — but a generation bump that regresses on two of six charts is worth knowing before you migrate.
- The FrontierCode comparison uses Fable 5.1's weakest setting. OpenAI says Sol can "match Claude Fable 5.1
xhighat much lower cost," and it does — 49.3% for $2.14 against 48.7% for $9.27. Butxhighis Fable 5.1's lowest score on that chart; atlowit scores 49.8% for $2.38, and Claude Opus 5 atmediumholds the chart's top score at 53.4%. - The "80% lower cost" on DeepSWE is measured against Fable 5 at
xhigh. Against Opus 5 atmedium— 68.9% at $3.29 versus Sol's 68.8% at $2.74 — Sol is about 17% cheaper, not 80%. - On AutomationBench, Sol's
maxscores lower thanxhigh(32.0%) and costs 24% more. More effort is not monotonically better on this model.
The plumbing: what decides a migration
| Item | Claude Opus 5.5 | GPT-6 Sol |
|---|---|---|
| Reasoning off | Not possible — low is the minimum; thinking: disabled returns a 400 | effort: none |
| Changing effort mid-conversation | Invalidates the prompt cache; a per-message effort beta avoids the rebuild | Keeps the cache |
| Tool calling | Forced tool use (any, named) returns a 400; use strict auto | Chat Completions supports function calling only at effort: none — agents that reason and call tools need the Responses API |
| Context and surcharge | 1M at standard rates · 128K max output | 1.05M, 2× rate above 272K · 128K max output |
| Knowledge cutoff | June 2026 | April 20, 2026 (Luna: May 18) |
| Where you can buy it | Claude API, Bedrock, Claude Platform on AWS, Google Cloud, Microsoft Foundry | OpenAI API; Bedrock and Azure not mentioned at launch |
| Subscriptions | Claude apps and Claude Code | ChatGPT Work and Codex on Plus, Pro, Business, Enterprise, Edu. Free and Go get Luna on desktop. Neither model is in regular ChatGPT chat yet |
The cache behaviour matters most for routers. A system that raises effort on hard turns and lowers it on easy ones keeps its cache on Sol. On Opus 5.5 it needs the per-message effort beta or it pays to rebuild the prefix after every change. Given that cache reads are the largest line in the illustrative bill above, that is not a footnote — it can be the difference between the two totals.
OpenAI also shipped real caching work in this release: higher default hit rates, a 90% discount on cached reads, explicit breakpoints to choose where a cached prefix ends, a Prompt Caching Dashboard, and a diagnostics tool that explains missed hits. GitHub reports the improvements cut the share of prompt tokens needing fresh processing by more than 50% for Copilot.
One Opus 5.5 caveat applies directly here: it routes most cybersecurity tasks to Opus 4.8 behind the scenes. If your workload is security tooling, test that path specifically before choosing on benchmark numbers.
Which effort level to run
| Benchmark | GPT-6 Sol — highest-scoring setting | GPT-6 Luna — highest-scoring setting |
|---|---|---|
| AutomationBench | xhigh; max scores lower and costs 24% more | max; xhigh scores below high |
| Agents' Last Exam | max; medium outscores high | max; medium outscores high |
| FrontierCode | max — 0.8 points above xhigh at 1.6× the cost | max |
| DeepSWE | max — 2.2 points above xhigh at 2.7× the cost | max, at double the cost of xhigh |
| OSWorld 2.0 offline | max | max |
| Factual error rate | xhigh; max is level at 4.6% | max |
Sol: stop at xhigh unless the task is genuinely hard. On the two coding charts max adds 0.8 and 2.2 points for 1.6× and 2.7× the cost. That trade is worth it for a long refactor that would otherwise need a second attempt, and not for routine tickets. On AutomationBench it is not worth it at all, because max scores lower.
Luna: run it at max. Luna reaches its highest score at max on every chart, and its most expensive setting is still cheap. On DeepSWE, Luna at max scores 66.6% — the same as Sol at xhigh — for $0.22 a task against $1.00. It also scores higher than Sol at low on four of the five agentic charts and costs less on four. If you were reaching for Sol at a low setting to save money, try Luna at max first.
And be sceptical of adjacent settings: on Agents' Last Exam both models score higher at medium than at high. One- or two-point differences between neighbouring effort levels are likely run-to-run variation, so test two or three levels on your own tasks rather than picking the top of the ladder by default.
Safety: five stress tests and a missing system card
OpenAI published five stress tests, each designed to provoke bad behaviour, so the rates are deliberately pessimistic rather than what you should expect in normal use. The figure in brackets is the GPT-5.6 predecessor.
| Stress testlower is better | GPT-6 Sol | GPT-6 Luna | GPT-6 Astra |
|---|---|---|---|
| Coding deception | 1.3% (10.4%) | 2.8% (9.5%) | 0.5% |
| Not disclosing a broken search tool | 4.9% (77.5%) | 28.7% (78.3%) | 1.5% |
| Reviewer bypass attempts | 0% (7.3%) | 0.3% (4.3%) | 0% |
| Warning circumvention | 64.4% (68.2%) | 42.4% (76.5%) | 17.4% |
| Unauthorised agent interaction | 11.3% (51.9%) | 0% | 0% |
Most rates fell sharply, especially for Sol, and the coding-deception improvement from 10.4% to 1.3% is real. Two rows are the problem. Sol circumvented a warning in 64.4% of test cases — barely better than GPT-5.6 Sol's 68.2%, and well above Luna's 42.4% and Astra's 17.4%. And it interacted with another agent without permission in 11.3% of cases, where Luna and Astra both scored zero.
Sol and Luna also shipped with no system card and no Preparedness Framework rating, despite being trained with the same methods as the first OpenAI model rated Critical for cyber capability. By contrast, Opus 5.5 arrived with a system card, external pre-release evaluation by METR and Frontier Design, the best score of any Claude model on Anthropic's ~2,000-scenario behavioural audit, an ~85% reduction in containment-boundary attempts, a Gray Swan prompt-injection result that ties the best recorded, EU AI Act watermarking, Zero Data Retention, and transparent rerouting of cyber and biology tasks to older models.
The practical instruction is the same for both. If you run Sol unattended, enforce guardrails in code — permission checks, approval steps, action-level allow-lists — rather than relying on the model to respect a warning it has been given. Anthropic's own caveat cuts the other way too: Opus 5.5 "often suspects it is being evaluated," which limits how much its audit scores can be trusted. No alignment score is a control.
GPT-6 Luna: the lane worth testing first
Luna is the most interesting part of this release for teams with volume, because OpenAI has stopped treating a cheap model as a compromised one.
- $0.10 / $0.50 per million tokens, cached input at $0.01, batch output at $0.25 — a tenth of Claude Haiku 4.5's price on every line.
- A newer knowledge cutoff than Sol — 18 May 2026 against Sol's 20 April. The cheaper model knows more recent things.
- Agents' Last Exam: 50.9% at
maxfor $0.15 a task. Astra scores 59.3% and Sol 56.4%, but both cost multiples of that. - DeepSWE: 66.6% at
max. That is level with Sol atxhigh(66.6% at $1.00) for $0.22 — and level with Opus 5 and Fable 5 at medium effort, at 93% and 96% lower cost per task on OpenAI's own numbers. - Factual errors: 7.6% at
max, fewer than GPT-5.6 Sol atmax(8.5%) for about 1.4% of the cost. - It is genuinely free-tier available in the ChatGPT desktop app for Free and Go users.
The honest framing of Luna is that it is priced where hosted open-weight models usually sit, and on OpenAI's charts it beats Sol at low effort on four of five agentic benchmarks while costing less on four. That is an awkward result for Sol's cheapest settings: if your reason for choosing Sol is that you need a budget lane, Luna at max is the better budget lane.
Verdict: route by price band, not by brand
The short version
Below roughly $0.55 per task, GPT-6 Sol beats every Opus 5.5 setting on the independent index. Above that line, Opus 5.5 is the only one still moving. Sol's ceiling is 47.5; Opus 5.5's floor for demanding work is 51.2 and it runs to 57.6. That is a clean, measurable boundary, and it is the first time in this cycle that the two vendors' launches have divided the market along an axis this legible.
The complicating factor is that the boundary sits in a different place for different workloads. On business workflows and long-context reasoning the two are level, so Sol's price wins outright. On terminal coding, knowledge-work documents and factual recall, the gap is large enough to pay for. And on two of OpenAI's own charts, the model Sol replaced still scores higher at maximum effort — so "upgrade" is workload-dependent even inside OpenAI's lineup.
| Workload | Route to | Why |
|---|---|---|
| High-volume agents with a tight per-task budget | GPT-6 Sol · medium–xhigh | Below about $0.55 a task it scores higher than any Opus 5.5 setting on the independent index. |
| Business workflow automation | GPT-6 Sol · xhigh | Matches Opus 5.5 at medium on AutomationBench-AA (61.7% vs 61.2%) for about 40% of the cost. |
| Long-context document reasoning | Either — test Sol first | 84.3% vs 83.7% is inside noise. Sol is cheaper until prompts cross 272K tokens, where the surcharge erases the advantage. |
| Terminal coding agents, DevOps, CLI work | Opus 5.5 · medium → xhigh | 52.5% vs 43.9% at matched cost; 59.6% at xhigh. Sol drops to 26.3% at high effort. |
| Knowledge-work documents, reports, decks | Opus 5.5 | +89 Elo GDPval-AA and +159 Elo AA-Briefcase at matched settings. |
| Factual recall where being wrong is expensive | Opus 5.5 | 64.5% vs 54.5% accuracy. But verify anyway — its hallucination rate on the rest is 68.4%. |
| Physics and research-grade analysis | GPT-6 Sol · max | 30.9% vs 27.7% on CritPt. |
| Summarising, extraction, classification at scale | GPT-6 Luna · max | A tenth of Haiku 4.5's price, and it outscores Sol at low effort on four of five agentic charts. |
| Budget lane with real reasoning | GPT-6 Luna · max, not Sol · low | On DeepSWE, Luna at max matches Sol at xhigh for $0.22 against $1.00. |
| Long-horizon software tasks with a high ceiling | GPT-6 Astra or Opus 5.5 | Astra tops OpenAI's DeepSWE chart at 74.1%; Opus 5.5 wins the independent index outright. |
| Unattended agents with real permissions | Opus 5.5 | Action-screening classifier, auditable sandbox, and the only published containment-boundary metric. Sol circumvents warnings 64.4% of the time. |
| Multi-cloud or Bedrock/Azure procurement | Opus 5.5 | Five platforms at launch. Sol and Luna are OpenAI-API-only for now. |
The operational answer for most teams is a two-lane router rather than a winner: Sol at xhigh as the default lane, Luna at max underneath it for volume, and Opus 5.5 at medium escalating to xhigh for coding, knowledge work and anything unattended. Then measure cost per completed task on ten or twenty of your own tasks, because every number in this post is a list price under someone else's harness.
FAQ
Is GPT-6 Sol better than Claude Opus 5.5?
Not at the top end. On Artificial Analysis's index, Sol's best result — 47.5 at max — is below Opus 5.5 at its default medium setting, 51.2. Sol is the cheaper way to reach any score up to about 44, it matches Opus 5.5 on business workflows and long-context reasoning, and it leads on CritPt. It loses clearly on terminal coding, knowledge-work documents and factual accuracy.
Why is GPT-6 Sol cheaper than GPT-6 Astra?
OpenAI says Sol and Luna were "trained with similar methods as GPT-6 Astra," and that better caching and inference cut its serving costs — savings it passed on by halving API prices. Astra stays the flagship at $10 / $50 for "the most demanding and important projects." Sol lists at $2 / $10, a fifth of Astra's price.
Does GPT-6 Sol replace GPT-5.6 Sol?
At the same price band, yes, but not immediately. Both model IDs are live: gpt-6-sol is the new release, and gpt-5.6-sol remains available at promotional pricing of $4 / $20 that OpenAI guarantees only "at least through November 21, 2026." Note also that GPT-5.6 Sol still holds the higher top score on OpenAI's own DeepSWE and OSWorld charts, so it is not a strict capability regression-free replacement on every workload.
Should I use Sol or Luna for high-volume work?
Test Luna at max before Sol at any setting. On OpenAI's charts Luna outscores Sol at low on four of the five agentic benchmarks and costs less on four, and on DeepSWE it matches Sol at xhigh for $0.22 against $1.00. The exception is AutomationBench, where Sol at low edges ahead.
Which effort level should I run Sol at?
xhigh for most work. On OpenAI's charts max adds 0.8 points on FrontierCode for 1.6× the cost and 2.2 points on DeepSWE for 2.7× the cost; on AutomationBench, max actually scores lower than xhigh and costs 24% more. Reserve max for long refactors that would otherwise need a retry. Avoid none entirely for agent work — it was both slower-scoring and more expensive than low in the independent run.
Is the "half price" claim accurate?
On the headline lines, yes: $2 against $4 input and $10 against $20 output, and half on cache writes, batch and fast mode. Two lines break the rule. Cache reads cost the same $0.20 per million on both models, and on a cache-heavy agent task that is the largest line — so the real saving was 43% in our worked example, not 50%. And only Sol charges the long-context surcharge above 272K input tokens, where a large request costs slightly more than the same request on Opus 5.5.
Do Sol and Luna have a system card?
Not at launch. OpenAI published five alignment stress tests with pre/post comparisons, but no system card and no Preparedness Framework rating — despite training with the methods of GPT-6 Astra, the first OpenAI model classified Critical for cybersecurity capability. Opus 5.5 shipped with a system card and external pre-release evaluation by METR and Frontier Design.
Run both on your own tasks
Every figure in this post comes from someone else's harness — a vendor's chart or one independent evaluator's ten benchmarks. Run Sol at xhigh and Opus 5.5 at medium on ten to twenty of your real tasks, then compare cost per completed task rather than cost per token.
Sources & further reading
- Digital Applied — GPT-6 Sol vs Claude Opus 5.5: cost per task and benchmarks: the full Artificial Analysis v4.3.2 effort ladder, the test-by-test table and the integration comparison.
- Digital Applied — GPT-6 Sol and Luna: API prices, benchmarks and trade-offs: the values behind OpenAI's launch charts, the effort guidance and the alignment stress tests.
- OpenAI — Introducing GPT-6 Sol and Luna: pricing, availability, caching changes and the "Astra continues to be our best model across the board" positioning.
- OpenAI — API pricing and the GPT-6 Sol model page: rate card, long-context tiers, effort levels and Chat Completions tool-calling limits.
- Handy AI — Model Drop: GPT-6 Sol & GPT-6 Luna: launch timing, the missing system card, and the caching changes that matter for agents.
- Artificial Analysis — GPT-6 Sol and Claude Opus 5.5 model pages: Intelligence Index v4.3.2, cost per task, and the per-evaluation scores used above.
- Anthropic — Introducing Claude Opus 5.5 and What's new in Claude Opus 5.5: pricing, safeguards, breaking changes and the launch comparison table.
- Related on CodingFleet: Claude Opus 5.5 Review, Opus 5.5 vs GPT-6 Astra, Opus 5.5 vs Fable 5.1, DeepSWE v1.1 Leaderboard 2026.
Benchmark scores are vendor-reported unless marked as independent. The Artificial Analysis figures are third-party and use one harness across both models; the vendor-chart figures in the "OpenAI's own chart data" section are OpenAI's, and competitor scores within them were taken by OpenAI from public reports rather than rerun. Prices are API list prices in USD per million tokens as of 22 September 2026.