OpenAI shipped GPT-6.1 Sol at DevDay on 29 September 2026, seven days after GPT-6 Sol and eight days after Anthropic's Claude Opus 5.5. The pitch is "near-Astra intelligence for a fifth of the price," but the more useful comparison is against the model it is actually priced against — because at $2 / $10 it is exactly half of Opus 5.5's $4 / $20, and Artificial Analysis has run both models across all ten effort levels in a single harness.
That independent run is what this article is built on, because the two vendors benchmarked against different things. OpenAI's charts compare Sol against Opus 5.5 on four evaluations it chose; Anthropic's launch table does not include GPT-6.1 Sol at all. The one harness that ran both on the same ten evaluations at every effort setting points somewhere neither lab emphasised.
TL;DR
GPT-6.1 Sol is a genuine substitute for Claude Opus 5.5 up to index 51 — and there is nothing above it. On Artificial Analysis Intelligence Index v4.3.2, Sol at xhigh scores exactly what Opus 5.5 scores at its default medium — 51 — for 71% less per task. Sol's ceiling is 52. Opus 5.5's default is 51. Every Opus setting above default (54, 56, 58) has no Sol counterpart at any price.
- At the ceilings: Opus 5.5 reaches 58 for $5.98 a task; Sol reaches 52 for $0.72. Six index points for 8.3× the cost.
- At the defaults: Opus 5.5
mediumis 51 for $1.34; Solmaxis 52 for $0.72. Sol is ahead and 46% cheaper. - At the floors: both score 42 at
low— Opus for $0.55, Sol for $0.13. A 76% saving for the same score. - The trap: past 272K input tokens, Sol's long-context rates apply to the entire request. Adding 1,000 tokens takes a request from $0.744 to $1.392 — +87%. Opus serves the full 1M at flat rates.
- The trade: Opus 5.5 wins 13 of 16 component evaluations. Sol wins GDP.pdf outright and ties CritPt.
where both models tie
$0.39 vs $1.34
above the crossover
272K input tokens
Two releases, eight days apart
The timing matters because three models landed inside a week. Claude Opus 5.5 shipped on 22 September, Claude Sonnet 5.5 on 28 September, and GPT-6.1 Sol on 29 September at OpenAI's DevDay. Sol's list price — $2 in, $10 out — is identical to Sonnet 5.5's, which makes the mid-tier the contested ground and Opus 5.5 the model both are undercutting.
GPT-6.1 Sol is not a new architecture. It is a point release that keeps GPT-6 Sol's standard rates, halves the cached-input rate to $0.10, and buys four points of Intelligence Index. Artificial Analysis's summary of the replacement is pointed: "GPT-6.1 Sol replaces GPT-6 Sol after just 7 days."
| Specification | Claude Opus 5.5 | GPT-6.1 Sol | GPT-6 Sol |
|---|---|---|---|
| Released | 22 September 2026 | 29 September 2026 | 22 September 2026 |
| API model ID | claude-opus-5-5 | gpt-6.1-sol | gpt-6-sol |
| Input | $4.00 / MTok | $2.00 / MTok | $2.00 / MTok |
| Cached input (read) | $0.20 / MTok | $0.10 / MTok | $0.20 / MTok |
| Cache write | $5.00 (5 min) · $8.00 (1 h) | $2.50 | $2.50 |
| Output | $20.00 / MTok | $10.00 / MTok | $10.00 / MTok |
| Above 272K input | no surcharge — 1M at flat rates | $4 / $0.20 / $15, whole request | same structure |
| Batch and Flex | $2 / $10 | $1 / $5 | $1 / $5 |
| Fast mode | $8 / $40 | $4 / $20 · Ultrafast 8× announced | $4 / $20 |
| Context window | 1,000,000 | 1,050,000 | 1,050,000 |
| Max output | 128K · 300K on Batches (beta) | 128K | 128K |
| Knowledge cutoff | June 2026 | April 30, 2026 | April 20, 2026 |
| Thinking control | Adaptive, always on — cannot be disabled | effort: low · medium · high · xhigh · max | "none" plus five levels |
| Default effort | medium | medium | medium |
| Cache contract | 512-token minimum · 5 min or 1 h retention | 1,024-token minimum · ≥30 min, refreshed on reuse | same as 6.1 |
| Where it runs | Claude API, Bedrock, Google Cloud, Microsoft Foundry, Claude Platform on AWS | OpenAI API, ChatGPT Work, Codex, Amazon Bedrock, OpenRouter — not in regular ChatGPT chat | same |
| System card | Opus 5.5 card published | addendum published; Critical cyber, High bio/chem | addendum only |
| Reasoning mode | effort only | effort plus reasoning.mode: "pro" in Responses | effort only |
One correction worth making early. Several third-party trackers still list Opus 5.5 at $5 / $25 — that is Opus 5's card. Anthropic's own pricing page and model documentation put Opus 5.5 at $4 input, $0.20 cached input, $20 output, with cache writes at $5 for five minutes or $8 for an hour. Every calculation below uses the published Opus 5.5 rates.
The ladder: all ten settings, one harness
This is the table the comparison should be built on. Both models ran the identical ten evaluations under Artificial Analysis — AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR v1.1 — and cost is measured per task including cache reads and writes.
| Configuration | Index | Cost / task | Output speed | Time per task |
|---|---|---|---|---|
GPT-6.1 Sol · low | 42 | $0.13 | 49 t/s | 80 s |
GPT-6.1 Sol · medium | 48 | $0.21 | 52 t/s | 154 s |
GPT-6.1 Sol · high | 50 | $0.32 | 53 t/s | 250 s |
GPT-6.1 Sol · xhigh | 51 | $0.39 | 59 t/s | 307 s |
Claude Opus 5.5 · low | 42 | $0.55 | 79 t/s | 79 s |
GPT-6.1 Sol · max | 52 | $0.72 | 58 t/s | 647 s |
Claude Opus 5.5 · medium default | 51 | $1.34 | 79 t/s | 205 s |
Claude Opus 5.5 · high | 54 | $1.82 | 79 t/s | 280 s |
Claude Opus 5.5 · xhigh | 56 | $3.46 | 80 t/s | 510 s |
Claude Opus 5.5 · max | 58 | $5.98 | 97 t/s | 783 s |
xhigh and Opus 5.5 at medium both score 51, and Sol does it for $0.39 against $1.34 — 71% cheaper. Source: Artificial Analysis Intelligence Index v4.3.2.Where the ladders cross, and where they stop
Read the ten rows in cost order and three findings fall out of the table without any interpretation.
1. At the floors, the scores are identical
Both models score 42 at low effort. Opus charges $0.55 a task for it; Sol charges $0.13. That is the cleanest single cost comparison in the release, because there is no capability argument to have — the scores match exactly and the price differs by 76%.
2. At the defaults, Sol is ahead on both axes
Opus 5.5's default is medium: 51 for $1.34. GPT-6.1 Sol's max setting is 52 for $0.72 — a point higher for 46% less. Sol's own default, medium, is 48 for $0.21, which is a different conversation: 84% cheaper than Opus with a three-point capability gap. The practical read is that Sol's top setting beats Opus's middle setting on both score and price.
3. Above index 51, Sol has nothing
This is the finding the vendor charts miss. Opus 5.5 at high, xhigh and max scores 54, 56 and 58. Sol's ceiling is 52, reached only at max, and it costs $0.72 to get there. Paying Sol more does not buy more: the four-step climb from xhigh to max costs 85% more per task for a single index point.
So the routing question is not "which is better." It is "does the task need an index score above 51?" If no, Sol is the cheaper answer at every level. If yes, Opus 5.5 is the only answer, and the premium starts at $1.82 a task.
A note on the numbers moving. Artificial Analysis has revised these figures since the GPT-6 Sol launch — Opus 5.5's ceiling and default both now read one point higher than in the earlier snapshot, and output speeds have been re-measured. The index is a live product; re-check the model pages before making a routing commitment on a single point of difference.
Test by test: sixteen components
The composite hides a lot. At maximum effort on both models, here is every component score Artificial Analysis publishes, with the leader marked.
| Evaluation | Opus 5.5 · max | GPT-6.1 Sol · max | Winner |
|---|---|---|---|
| Economics | 66 | 59 | Opus 5.5 |
| Strategy & Ops | 64 | 57 | Opus 5.5 |
| Legal | 63 | 58 | Opus 5.5 |
| Finance & Accounting | 61 | 54 | Opus 5.5 |
| Healthcare & Medical | 61 | 51 | Opus 5.5 |
| Engineering | 60 | 54 | Opus 5.5 |
| AA-Briefcase v1.1 | 1,807 | 1,564 | Opus 5.5 · +243 Elo |
| GDPval-AA v2.1 | 1,866 | 1,575 | Opus 5.5 · +291 Elo |
| AutomationBench-AA | 69.5% | 64.9% | Opus 5.5 |
| Terminal-Bench 4.0 | 59.6% | 56.1% | Opus 5.5 |
| SciCode | 66.9% | 54.2% | Opus 5.5 · +12.7 |
| Humanity's Last Exam | 61.4% | 52.9% | Opus 5.5 · +8.5 |
| AA-LCR v1.1 long-context reasoning | 84.7% | 83.0% | Opus 5.5 |
| AA-Omniscience | 46 | 42 | Opus 5.5 |
| GDP.pdf complex professional documents | 26.2% | 31.0% | GPT-6.1 Sol · +4.8 |
| CritPt physics research | 31.7% | 31.7% | tie |
Opus 5.5 wins thirteen of sixteen and ties one. Its margins are widest exactly where you would expect a $4/$20 frontier-class model to win: knowledge-work documents (+291 Elo on GDPval-AA, +243 on AA-Briefcase), scientific coding (+12.7 on SciCode) and expert reasoning (+8.5 on Humanity's Last Exam).
Sol's one clear win is GDP.pdf, the Surge AI benchmark over dense professional PDFs — financial statements, legal filings, multi-column tables. Sol scores 31.0% at max and 32.0% at high, against Opus 5.5's 26.2% at max. OpenAI's own chart puts its figure at 32.0% against Opus 5.5 at 28.8% "with fallbacks," and against GPT-6 Astra's 32.2%. For document-extraction pipelines, that is a real and specific advantage — and it is the only one.
An oddity in both models: more effort is not always better
GDP.pdf peaks in the middle of the ladder for both models. Opus 5.5 is strongest at high (28.8%) and weaker at max (26.2%). Sol is strongest at high (32.0%) and weaker at max (31.0%). And on AutomationBench-AA, Sol's best result is at xhigh (66.6%), above its own max (64.9%).
OpenAI's own coding chart shows the same shape more dramatically, which is worth stating precisely because it determines the setting you should ship:
high and declines. On OpenAI's own DeepSWE v1.1 chart, GPT-6.1 Sol scores 75.2% at high for $0.65 a task, then falls to 71.9% at xhigh ($0.79) and max ($1.57). At high it beats GPT-6 Astra's best run of 74.8% for roughly an eighth of Astra's $3.92–$7.50. Source: OpenAI launch chart data.For Sol, the operational instruction is unambiguous: ship high for coding, medium for business workflows, and treat max as unproven on your workload. The same warning applies with less force to Opus 5.5, where GDP.pdf regresses at max even as the composite index keeps climbing.
The 272K cliff
GPT-6.1 Sol's price advantage is usually quoted as 50%. That is true for prompts up to 272,000 input tokens. Past that threshold, OpenAI applies its long-context rates to the whole request — not just the portion above the line. Anthropic serves Opus 5.5's full 1M-token window at flat rates.
The difference is not gradual. Here is a request with 20,000 output tokens at both sides of the threshold:
That threshold turns ordinary agent decisions into cost decisions. An agent that keeps appending full documents to its context, or returns large tool results without compaction, will cross 272K without anyone intending it to. But the reverse is also true, and worth saying: compaction and retrieval discipline are now worth real money on Sol specifically, because staying under the line keeps the 50% advantage intact.
Where the half-price headline holds, and where it evaporates
| Reproducible usage | GPT-6.1 Sol | Claude Opus 5.5 | Sol's saving |
|---|---|---|---|
| 10K fresh input + 2K output | $0.04 | $0.08 | 50% |
| 100K fresh input + 20K output | $0.40 | $0.80 | 50% |
| 100K cache read + 10K input + 10K output | $0.13 | $0.26 | 50% |
| 20 requests: one 100K prefix write, 19 cache reads, 5K output each | $1.44 | $2.88 | 50% |
| 400K fresh input + 20K output over the line | $1.90 | $2.00 | 5% |
| 900K fresh input + 10K output over the line | $3.75 | $3.80 | 1.3% |
The pattern is clean: short prompts and output-heavy work keep the full 50%. Large fresh document inputs with little output — the shape of RAG pipelines, contract review and compliance screening — collapse it to between 1% and 5%. If your workload looks like the last two rows, the price is not the reason to switch.
The caching contracts are not the same product. Sol requires a 1,024-token minimum and retains a cache for at least 30 minutes from the last write or reuse, refreshing on reuse. Opus 5.5's minimum is 512 tokens with a five-minute default or a pricier one-hour option. Sol's cheaper cache reads are real, but a workflow that goes quiet for an hour, or whose prefix shifts between calls, may find Opus's longer retention cheaper in practice despite the higher rate.
Speed: the counterintuitive part
Both vendors market throughput, and both are right about their own number. What the independent run shows is that throughput and task time are different quantities, and they favour different models.
| At matched cost band | Output speed | Time to first token | Time per index task |
|---|---|---|---|
GPT-6.1 Sol · medium — $0.21 | 52 t/s | 5.4 s | 154 s |
Claude Opus 5.5 · medium — $1.34 | 79 t/s | 23.5 s | 205 s |
GPT-6.1 Sol · max — $0.72 | 58 t/s | 279 s | 647 s |
Claude Opus 5.5 · max — $5.98 | 97 t/s | 745 s | 783 s |
Opus 5.5 streams faster at every setting — 97 tokens per second at max against Sol's 58 — which is why Anthropic's "more than 30% faster output" claim holds. But it takes 12.4 minutes to produce its first token at max effort, against 4.7 minutes for Sol. End to end, Sol finishes an index task in 647 seconds where Opus takes 783.
That has a practical consequence for interactive use. A model's streaming speed is what you feel while reading an answer; its time-to-first-token is what you feel while waiting for one to start. For a chat interface, Opus 5.5's 5.4-second-equivalent settings matter — at medium it starts in 23.5 seconds against Sol's 5.4. For a batch agent that runs unattended, only the total time per task counts, and Sol wins there at every comparable price point.
The same model, three different scores
The most instructive part of this comparison is how much Opus 5.5's numbers move depending on who ran them. These are not errors; they are different harnesses, and they should put a floor under how much weight you give any single figure.
| Evaluation | Anthropic's own table | OpenAI's launch run | Independent |
|---|---|---|---|
| GDP.pdf | not published | 28.8% with fallbacks | 26.2% at max, 28.8% at high |
| Terminal-Bench-Science 0.1 | 58.7% | 63.3% | not in the index |
| AutomationBench | 40.0% at max (Zapier) | 42.5% at max | 69.5% at max (partial credit) |
| OSWorld | 81.8% on 2.0 partial | 60.3% on 2.0 offline | not in the index |
| Intelligence Index v4.3.2 | not participated | not participated | 58 at max, 51 at medium |
Three things follow. First, AutomationBench is two benchmarks sharing a name: Zapier's binary workflow score and Artificial Analysis's partial-credit version read 40.0% and 69.5% for the same model at the same effort. Quoting either without the qualifier is misleading. Second, OpenAI's GDP.pdf figure for Opus 5.5 excludes nothing but is labelled "with fallbacks", and OpenAI itself notes that its AutomationBench figure for Claude Fable 5.1 omits fallback cost that occurred on roughly 40% of tasks. Third, on Anthropic's side, the Opus 5.5 launch table has no row for GPT-6.1 Sol — it lists GPT-5.6 Sol, the model OpenAI replaced eight days before.
Neither vendor is being dishonest. Both are showing where their model looks best. The reason the Artificial Analysis table anchors this article is simply that it is the only run where both models faced the same tests under the same operator.
Safety and safeguards: two different bets
GPT-6.1 Sol
- Classified Critical for cybersecurity and High for biological and chemical capability under OpenAI's framework, carrying Astra's safeguard stack.
- ExploitBench 99.7% at max against GPT-6 Sol's 81.7% — with OpenAI's own caveat that historical-vulnerability contamination may inflate this.
- Broken-tool honesty improved: when a search or API tool was deliberately degraded, failure to tell the user fell from 4.9% to 2.1% (Astra: 1.5%).
- Zero bypass attempts in automated safety-reviewer and authorised-scope evaluations.
- Factual error rate at low effort fell from 11.4% to 7.7% on user-flagged historical errors, staying within 1.9 points of Astra at every level.
- Instruction following: OpenAI reports clear gains on negative constraints such as "do not modify files outside
/src".
Claude Opus 5.5
- Best score to date on Anthropic's automated behavioural audit — roughly 2,000 scenarios — with containment-boundary attempts about 85% lower than Opus 5, all low severity and self-reported.
- Transparent rerouting: cyber tasks go to Opus 4.8, biology and frontier-LLM tasks to Opus 5, with the switch visible rather than silent.
- Gray Swan prompt-injection benchmark: ties Fable 5.1 for the lowest success rate recorded.
- Vetted access scales through the Cyber Verification Program (three tiers, up to Mythos models) and the Life Sciences Verification Program.
- EU AI Act watermarking and Zero Data Retention support.
- However: Anthropic states Opus 5.5 "often suspects it is being evaluated," which it says limits the reliability of its own audit scores.
The fallback policy changes what the benchmarks measure. When Opus 5.5's classifiers fire, work is completed by a different model — and Vals records how often that happened on its index: 91 of 2,281 tasks (3.99%) were fallback-assisted, with 18 refusals (0.79%) scored as failures. On SRE-Bench, outside the index, 82.82% of Opus's tasks were fallback-assisted and 10.30% were refusals, which is why Opus scores 33.59% there against Sol's 50.76%. That is not Opus 5.5 doing security work badly; it is Opus 5.5 declining to do it and handing it to Opus 4.8.
Read positively, that is the safeguard working. Read as a benchmark, it means a "Claude Opus 5.5" row can be a mixture of three models. For security and biology workloads specifically, check your refusal path before comparing scores.
Operational changes to audit before you switch
effort: "none"is gone. GPT-6.1 Sol drops the non-reasoning setting that GPT-6 Sol offered. OpenAI's migration guide says to uselowinstead, so any caller settingnoneexplicitly needs a config change and a regression run.- Sol uses 10–30% more output tokens than GPT-6 Sol at each effort setting, per Artificial Analysis. The cost cut came from the rate card, not from efficiency.
- Sol is slower to stream: 58–59 tokens per second at its top settings against GPT-6 Sol's higher figures and an average of about 70 across tracked models. Latency-sensitive interfaces should be tested before migrating.
- Opus 5.5 cannot disable thinking at all. If your integration shuts thinking off to control cost, it now returns a 400 —
loweffort is the replacement. - Refusal billing changed on 24 September 2026. Anthropic resumed charging for pre-output refusals in the
bio,frontier_llmandreasoning_extractioncategories;cyberrefusals remain unbilled. If your pipeline triggers safeguards, that is a new line on the invoice — and a fallback is charged separately at the fallback model's rates. - Cache minimums differ: 512 tokens for Opus 5.5, 1,024 for Sol. Short system prompts that cache on one model may not cache on the other.
- Availability: Sol is in the API, ChatGPT Work, Codex, Bedrock and OpenRouter but not in regular ChatGPT chat, and it is off by default in Enterprise and Edu workspaces until an admin enables it. Opus 5.5 is on five cloud platforms from day one.
- One dead end to name: "cost per accepted task" is the right production metric, and neither vendor publishes it. If Sol attempts a task at $0.40 and passes 40% of the time while Opus attempts at $0.80 and passes 90%, the arithmetic is close to a wash despite a 2× sticker gap. Measure success rates on your own queue.
Verdict
The short version
GPT-6.1 Sol is the better default and Claude Opus 5.5 is the better ceiling, and the line between them is sharp enough to route on. Below index 51, Sol is cheaper at every single level — often by 70–76% for an identical score. Above index 51, Opus 5.5 is the only model that goes there, and the entry price is $1.82 a task, rising to $5.98 at the top.
The decision rule that follows is unusually simple: run Sol by default, and escalate to Opus 5.5 when a task fails a checkpoint — or when you know in advance it needs an index score above 51. Categories that reliably need the ceiling are knowledge-work documents, scientific coding, and open-ended expert reasoning, where Opus's margins are 8 to 13 points on the independent run. Categories that reliably do not are document extraction, high-volume business workflows, and anything short-prompt and output-heavy.
Two caveats before committing. Sol's half-price advantage is real only under 272K input tokens — on large-input, short-output workloads it shrinks to between 1% and 5%. And Sol's own effort ladder peaks at high on coding, so paying for max is likely to buy nothing.
| Your workload | Route to | Why |
|---|---|---|
| High-volume extraction, classification, routing | GPT-6.1 Sol | $0.13 a task at low for the same index score Opus charges $0.55 for |
| Complex professional PDFs and dense documents | GPT-6.1 Sol | The one component it wins outright: 31.0% vs 26.2% on GDP.pdf |
| Business workflow automation | GPT-6.1 Sol at medium | 35.4% against Opus at 29.5% on OpenAI's run, for about a third of the cost |
| Budget-capped agent loops with heavy caching | GPT-6.1 Sol | $0.10 cached input and the cheapest cost-per-task figure on the frontier |
| Latency-sensitive interactive agents | GPT-6.1 Sol at lower effort | 5.4 s to first token at medium against 23.5 s for Opus at the same setting |
| Terminal and CI coding agents | Either — test both | Terminal-Bench 4.0 reads 59.6% vs 56.1% on one harness and 65.15% vs 55.05% on another |
| Knowledge work: briefs, decks, analysis | Claude Opus 5.5 | +291 Elo GDPval-AA, +243 Elo AA-Briefcase |
| Scientific coding and physics | Claude Opus 5.5 | SciCode 66.9% vs 54.2%; CritPt is a tie at 31.7% |
| Expert-level reasoning and hard science agents | Claude Opus 5.5 at high+ | HLE 61.4% vs 52.9%; the only settings that reach index 54–58 |
| Workloads past 272K input on every request | Claude Opus 5.5 | Flat 1M pricing; Sol's advantage collapses to 1–5% |
| Authorised offensive security | Neither — different tools | Sol is Critical-rated and gated; Opus 5.5 reroutes the work to Opus 4.8 |
| Absolute cost floor | GPT-6.1 Sol at low | $0.13 a task; nothing cheaper sits at index 42 |
FAQ
Is GPT-6.1 Sol better than Claude Opus 5.5?
Depends entirely on the effort setting. At low both score 42 and Sol is 76% cheaper. At Opus's default medium (51, $1.34), Sol's xhigh matches the score for $0.39 and Sol's max beats it (52) for $0.72. Above index 51, Opus 5.5 has no competitor: it reaches 54, 56 and 58 where Sol stops at 52. So Sol is better below the crossover and Opus is the only option above it.
Why is Sol exactly half the price?
Input is $2 against $4, output is $10 against $20, and cache writes are $2.50 against $5. Cache reads are the exception where Sol is five times cheaper — $0.10 against $0.20 — because OpenAI cut them when it shipped GPT-6.1 Sol. Sol is not uniformly half price: it is half on most lines, a fifth on cached input, and then more expensive above 272K input because of the long-context surcharge that Opus does not apply.
What is the 272K cliff and does it affect me?
OpenAI applies its long-context rates to the entire request once input exceeds 272,000 tokens: $4 input, $0.20 cache read, $15 output. A request with 272K input and 20K output costs $0.744; at 273K the same request costs $1.392 — an 87% jump for a thousand extra tokens. If your prompts routinely exceed 272K, or your agent appends documents without compaction, assume the half-price advantage does not apply to you. Opus 5.5 serves the full 1M window at flat rates.
Should I use max effort on GPT-6.1 Sol?
Probably not. On OpenAI's DeepSWE v1.1 chart, Sol peaks at high (75.2% for $0.65) and falls to 71.9% at both xhigh and max, the latter costing $1.57 — 2.4× the price for 2.4 points less. On Artificial Analysis's index, the last step from xhigh to max buys one point for 85% more. Ship high for coding and medium for workflows, and test max on your own workload before paying for it.
Does Opus 5.5 cost $4/$20 or $5/$25?
$4 input and $20 output per million tokens, with cached input at $0.20 and cache writes at $5 for five minutes or $8 for an hour. Some third-party price trackers still display $5/$25, which is Claude Opus 5's card — Opus 5.5 cut prices when it shipped on 22 September 2026. Check Anthropic's own pricing page before quoting a number to anyone.
Which model is more honest about its own failures?
They are measured differently and both improved. OpenAI reports that when a search or API tool was deliberately broken, Sol failed to tell the user in 2.1% of runs, down from 4.9% for GPT-6 Sol and close to Astra's 1.5%. On Artificial Analysis's AA-Omniscience, Opus 5.5 scores 46 against Sol's 42 on a scale that rewards correct answers and penalises hallucinations without penalising abstention. Opus 5.5 remains the stronger model on calibration by the independent measure; Sol is the cheaper one.
Can I use both in one pipeline?
Yes, and the ladder makes the rule concrete. Route to Sol by default; escalate to Opus 5.5 when a task fails a checkpoint, or when you know in advance it needs an index score above 51. The two do not share prompt format, tool schemas or reasoning-transcript conventions, so treat the boundary as a full task handoff rather than a mid-conversation model switch — unlike the Opus 5.5 to Fable 5.1 handoff, which does preserve thinking blocks on the same API.
Run the crossover on your own queue
Take fifty tasks you have already solved, run them at Sol's high and Opus 5.5's medium, and count the failures. That single test will tell you whether your work sits above or below the index-51 line.
Sources & further reading
- Artificial Analysis — GPT-6.1 Sol vs Claude Opus 5.5 release comparison: all ten effort levels, index scores, cost per task, component evaluations, output speed and time per task.
- Artificial Analysis — GPT-6.1 Sol replaces GPT-6 Sol after just 7 days: the independent launch analysis, cost-efficiency frontier and token-usage findings.
- OpenAI — Introducing GPT-6.1 Sol: DeepSWE v1.1, OSWorld 2.0, GDP.pdf, AutomationBench, Terminal-Bench Science and factuality charts.
- OpenAI — GPT-6.1 Sol system card addendum: capability classifications, safety evaluations, broken-tool disclosure and guardrail adherence.
- Anthropic — Introducing Claude Opus 5.5 and the Opus 5.5 model documentation: pricing, benchmark table, context and output limits.
- Claude Platform Docs — Refusals and fallback and the release notes (24 September 2026): refusal billing categories and fallback credit.
- Vals — Index v2.1 with the GPT-6.1 Sol and Opus 5.5 model pages: the second independent composite, plus fallback and refusal accounting.
- Kingy AI — GPT-6.1 Sol vs Claude Opus 5.5: benchmarks, specs and cost per task: the reproducible per-request cost table and the 272K threshold arithmetic.
- Vellum — GPT-6.1 Sol benchmarks explained: chart-data tables for DeepSWE, OSWorld 2.0, GDP.pdf and AutomationBench.
- DataCamp — GPT-6.1 Sol: features, benchmarks, pricing and access.
- Related on CodingFleet: Claude Opus 5.5 Review, Claude Opus 5.5 vs GPT-6 Sol, GPT-6 Sol vs GPT-5.6 Sol, GPT-6 Astra Review.
Benchmark scores are vendor- or evaluator-reported and labelled as such throughout. Effort settings are stated wherever the source states them; where a vendor omits the effort level of a comparison, this article says so rather than assuming one. Cross-harness numbers are directional. Prices are USD per million tokens as of early October 2026.