OpenAI shipped GPT-6.1 Sol at DevDay on 29 September 2026, seven days after GPT-6 Sol and eight days after Anthropic's Claude Opus 5.5. The pitch is "near-Astra intelligence for a fifth of the price," but the more useful comparison is against the model it is actually priced against — because at $2 / $10 it is exactly half of Opus 5.5's $4 / $20, and Artificial Analysis has run both models across all ten effort levels in a single harness.

That independent run is what this article is built on, because the two vendors benchmarked against different things. OpenAI's charts compare Sol against Opus 5.5 on four evaluations it chose; Anthropic's launch table does not include GPT-6.1 Sol at all. The one harness that ran both on the same ten evaluations at every effort setting points somewhere neither lab emphasised.

TL;DR

GPT-6.1 Sol is a genuine substitute for Claude Opus 5.5 up to index 51 — and there is nothing above it. On Artificial Analysis Intelligence Index v4.3.2, Sol at xhigh scores exactly what Opus 5.5 scores at its default medium — 51 — for 71% less per task. Sol's ceiling is 52. Opus 5.5's default is 51. Every Opus setting above default (54, 56, 58) has no Sol counterpart at any price.

  • At the ceilings: Opus 5.5 reaches 58 for $5.98 a task; Sol reaches 52 for $0.72. Six index points for 8.3× the cost.
  • At the defaults: Opus 5.5 medium is 51 for $1.34; Sol max is 52 for $0.72. Sol is ahead and 46% cheaper.
  • At the floors: both score 42 at low — Opus for $0.55, Sol for $0.13. A 76% saving for the same score.
  • The trap: past 272K input tokens, Sol's long-context rates apply to the entire request. Adding 1,000 tokens takes a request from $0.744 to $1.392 — +87%. Opus serves the full 1M at flat rates.
  • The trade: Opus 5.5 wins 13 of 16 component evaluations. Sol wins GDP.pdf outright and ties CritPt.
51
the crossover index
where both models tie
−71%
Sol's cost at that score
$0.39 vs $1.34
58 vs 52
index ceilings
above the crossover
+87%
cost jump crossing
272K input tokens

Two releases, eight days apart

The timing matters because three models landed inside a week. Claude Opus 5.5 shipped on 22 September, Claude Sonnet 5.5 on 28 September, and GPT-6.1 Sol on 29 September at OpenAI's DevDay. Sol's list price — $2 in, $10 out — is identical to Sonnet 5.5's, which makes the mid-tier the contested ground and Opus 5.5 the model both are undercutting.

GPT-6.1 Sol is not a new architecture. It is a point release that keeps GPT-6 Sol's standard rates, halves the cached-input rate to $0.10, and buys four points of Intelligence Index. Artificial Analysis's summary of the replacement is pointed: "GPT-6.1 Sol replaces GPT-6 Sol after just 7 days."

SpecificationClaude Opus 5.5GPT-6.1 SolGPT-6 Sol
Released22 September 202629 September 202622 September 2026
API model IDclaude-opus-5-5gpt-6.1-solgpt-6-sol
Input$4.00 / MTok$2.00 / MTok$2.00 / MTok
Cached input (read)$0.20 / MTok$0.10 / MTok$0.20 / MTok
Cache write$5.00 (5 min) · $8.00 (1 h)$2.50$2.50
Output$20.00 / MTok$10.00 / MTok$10.00 / MTok
Above 272K inputno surcharge — 1M at flat rates$4 / $0.20 / $15, whole requestsame structure
Batch and Flex$2 / $10$1 / $5$1 / $5
Fast mode$8 / $40$4 / $20 · Ultrafast 8× announced$4 / $20
Context window1,000,0001,050,0001,050,000
Max output128K · 300K on Batches (beta)128K128K
Knowledge cutoffJune 2026April 30, 2026April 20, 2026
Thinking controlAdaptive, always on — cannot be disabledeffort: low · medium · high · xhigh · max"none" plus five levels
Default effortmediummediummedium
Cache contract512-token minimum · 5 min or 1 h retention1,024-token minimum · ≥30 min, refreshed on reusesame as 6.1
Where it runsClaude API, Bedrock, Google Cloud, Microsoft Foundry, Claude Platform on AWSOpenAI API, ChatGPT Work, Codex, Amazon Bedrock, OpenRouter — not in regular ChatGPT chatsame
System cardOpus 5.5 card publishedaddendum published; Critical cyber, High bio/chemaddendum only
Reasoning modeeffort onlyeffort plus reasoning.mode: "pro" in Responseseffort only

One correction worth making early. Several third-party trackers still list Opus 5.5 at $5 / $25 — that is Opus 5's card. Anthropic's own pricing page and model documentation put Opus 5.5 at $4 input, $0.20 cached input, $20 output, with cache writes at $5 for five minutes or $8 for an hour. Every calculation below uses the published Opus 5.5 rates.

The ladder: all ten settings, one harness

This is the table the comparison should be built on. Both models ran the identical ten evaluations under Artificial Analysis — AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR v1.1 — and cost is measured per task including cache reads and writes.

ConfigurationIndexCost / taskOutput speedTime per task
GPT-6.1 Sol · low42$0.1349 t/s80 s
GPT-6.1 Sol · medium48$0.2152 t/s154 s
GPT-6.1 Sol · high50$0.3253 t/s250 s
GPT-6.1 Sol · xhigh51$0.3959 t/s307 s
Claude Opus 5.5 · low42$0.5579 t/s79 s
GPT-6.1 Sol · max52$0.7258 t/s647 s
Claude Opus 5.5 · medium default51$1.3479 t/s205 s
Claude Opus 5.5 · high54$1.8279 t/s280 s
Claude Opus 5.5 · xhigh56$3.4680 t/s510 s
Claude Opus 5.5 · max58$5.9897 t/s783 s
GPT-6.1 Sol · low 42 $0.13 GPT-6.1 Sol · medium 48 $0.21 GPT-6.1 Sol · high 50 $0.32 GPT-6.1 Sol · xhigh 51 $0.39 Claude Opus 5.5 · low 42 $0.55 GPT-6.1 Sol · max 52 $0.72 Claude Opus 5.5 · medium 51 $1.34 Claude Opus 5.5 · high 54 $1.82 Claude Opus 5.5 · xhigh 56 $3.46 Claude Opus 5.5 · max 58 $5.98 Claude Opus 5.5 GPT-6.1 Sol Sorted by cost per Intelligence Index task · bar length = index score
Two clusters, one overlap. Everything green sits below $0.72 a task; everything coral starts at $0.55 and runs to $5.98. The single crossover is real and precise: Sol at xhigh and Opus 5.5 at medium both score 51, and Sol does it for $0.39 against $1.34 — 71% cheaper. Source: Artificial Analysis Intelligence Index v4.3.2.

Where the ladders cross, and where they stop

Read the ten rows in cost order and three findings fall out of the table without any interpretation.

1. At the floors, the scores are identical

Both models score 42 at low effort. Opus charges $0.55 a task for it; Sol charges $0.13. That is the cleanest single cost comparison in the release, because there is no capability argument to have — the scores match exactly and the price differs by 76%.

2. At the defaults, Sol is ahead on both axes

Opus 5.5's default is medium: 51 for $1.34. GPT-6.1 Sol's max setting is 52 for $0.72 — a point higher for 46% less. Sol's own default, medium, is 48 for $0.21, which is a different conversation: 84% cheaper than Opus with a three-point capability gap. The practical read is that Sol's top setting beats Opus's middle setting on both score and price.

3. Above index 51, Sol has nothing

This is the finding the vendor charts miss. Opus 5.5 at high, xhigh and max scores 54, 56 and 58. Sol's ceiling is 52, reached only at max, and it costs $0.72 to get there. Paying Sol more does not buy more: the four-step climb from xhigh to max costs 85% more per task for a single index point.

So the routing question is not "which is better." It is "does the task need an index score above 51?" If no, Sol is the cheaper answer at every level. If yes, Opus 5.5 is the only answer, and the premium starts at $1.82 a task.

A note on the numbers moving. Artificial Analysis has revised these figures since the GPT-6 Sol launch — Opus 5.5's ceiling and default both now read one point higher than in the earlier snapshot, and output speeds have been re-measured. The index is a live product; re-check the model pages before making a routing commitment on a single point of difference.

Test by test: sixteen components

The composite hides a lot. At maximum effort on both models, here is every component score Artificial Analysis publishes, with the leader marked.

EvaluationOpus 5.5 · maxGPT-6.1 Sol · maxWinner
Economics6659Opus 5.5
Strategy & Ops6457Opus 5.5
Legal6358Opus 5.5
Finance & Accounting6154Opus 5.5
Healthcare & Medical6151Opus 5.5
Engineering6054Opus 5.5
AA-Briefcase v1.11,8071,564Opus 5.5 · +243 Elo
GDPval-AA v2.11,8661,575Opus 5.5 · +291 Elo
AutomationBench-AA69.5%64.9%Opus 5.5
Terminal-Bench 4.059.6%56.1%Opus 5.5
SciCode66.9%54.2%Opus 5.5 · +12.7
Humanity's Last Exam61.4%52.9%Opus 5.5 · +8.5
AA-LCR v1.1 long-context reasoning84.7%83.0%Opus 5.5
AA-Omniscience4642Opus 5.5
GDP.pdf complex professional documents26.2%31.0%GPT-6.1 Sol · +4.8
CritPt physics research31.7%31.7%tie

Opus 5.5 wins thirteen of sixteen and ties one. Its margins are widest exactly where you would expect a $4/$20 frontier-class model to win: knowledge-work documents (+291 Elo on GDPval-AA, +243 on AA-Briefcase), scientific coding (+12.7 on SciCode) and expert reasoning (+8.5 on Humanity's Last Exam).

Sol's one clear win is GDP.pdf, the Surge AI benchmark over dense professional PDFs — financial statements, legal filings, multi-column tables. Sol scores 31.0% at max and 32.0% at high, against Opus 5.5's 26.2% at max. OpenAI's own chart puts its figure at 32.0% against Opus 5.5 at 28.8% "with fallbacks," and against GPT-6 Astra's 32.2%. For document-extraction pipelines, that is a real and specific advantage — and it is the only one.

An oddity in both models: more effort is not always better

GDP.pdf peaks in the middle of the ladder for both models. Opus 5.5 is strongest at high (28.8%) and weaker at max (26.2%). Sol is strongest at high (32.0%) and weaker at max (31.0%). And on AutomationBench-AA, Sol's best result is at xhigh (66.6%), above its own max (64.9%).

OpenAI's own coding chart shows the same shape more dramatically, which is worth stating precisely because it determines the setting you should ship:

low · $0.17 64.4% medium · $0.42 73.0% high · $0.65 75.2% peak xhigh · $0.79 71.9% max · $1.57 71.9% 2.4 points lower, 2.4× the price
Sol's coding score peaks at high and declines. On OpenAI's own DeepSWE v1.1 chart, GPT-6.1 Sol scores 75.2% at high for $0.65 a task, then falls to 71.9% at xhigh ($0.79) and max ($1.57). At high it beats GPT-6 Astra's best run of 74.8% for roughly an eighth of Astra's $3.92–$7.50. Source: OpenAI launch chart data.

For Sol, the operational instruction is unambiguous: ship high for coding, medium for business workflows, and treat max as unproven on your workload. The same warning applies with less force to Opus 5.5, where GDP.pdf regresses at max even as the composite index keeps climbing.

The 272K cliff

GPT-6.1 Sol's price advantage is usually quoted as 50%. That is true for prompts up to 272,000 input tokens. Past that threshold, OpenAI applies its long-context rates to the whole request — not just the portion above the line. Anthropic serves Opus 5.5's full 1M-token window at flat rates.

The difference is not gradual. Here is a request with 20,000 output tokens at both sides of the threshold:

GPT-6.1 Sol · 272K in $0.744 GPT-6.1 Sol · 273K in $1.392 Claude Opus 5.5 · 272K in $1.488 Claude Opus 5.5 · 273K in $1.492
One thousand tokens costs Sol 87% more. Its bill moves from $0.744 to $1.392 while Opus 5.5's moves from $1.488 to $1.492 — a 0.3% change. A Sol request that crosses the line is still cheaper than Opus, but only by 7%, and only until output volume changes the ratio. Calculations from published rate schedules, 20K output tokens, no cache activity.

That threshold turns ordinary agent decisions into cost decisions. An agent that keeps appending full documents to its context, or returns large tool results without compaction, will cross 272K without anyone intending it to. But the reverse is also true, and worth saying: compaction and retrieval discipline are now worth real money on Sol specifically, because staying under the line keeps the 50% advantage intact.

Where the half-price headline holds, and where it evaporates

Reproducible usageGPT-6.1 SolClaude Opus 5.5Sol's saving
10K fresh input + 2K output$0.04$0.0850%
100K fresh input + 20K output$0.40$0.8050%
100K cache read + 10K input + 10K output$0.13$0.2650%
20 requests: one 100K prefix write, 19 cache reads, 5K output each$1.44$2.8850%
400K fresh input + 20K output over the line$1.90$2.005%
900K fresh input + 10K output over the line$3.75$3.801.3%

The pattern is clean: short prompts and output-heavy work keep the full 50%. Large fresh document inputs with little output — the shape of RAG pipelines, contract review and compliance screening — collapse it to between 1% and 5%. If your workload looks like the last two rows, the price is not the reason to switch.

The caching contracts are not the same product. Sol requires a 1,024-token minimum and retains a cache for at least 30 minutes from the last write or reuse, refreshing on reuse. Opus 5.5's minimum is 512 tokens with a five-minute default or a pricier one-hour option. Sol's cheaper cache reads are real, but a workflow that goes quiet for an hour, or whose prefix shifts between calls, may find Opus's longer retention cheaper in practice despite the higher rate.

Speed: the counterintuitive part

Both vendors market throughput, and both are right about their own number. What the independent run shows is that throughput and task time are different quantities, and they favour different models.

At matched cost bandOutput speedTime to first tokenTime per index task
GPT-6.1 Sol · medium — $0.2152 t/s5.4 s154 s
Claude Opus 5.5 · medium — $1.3479 t/s23.5 s205 s
GPT-6.1 Sol · max — $0.7258 t/s279 s647 s
Claude Opus 5.5 · max — $5.9897 t/s745 s783 s

Opus 5.5 streams faster at every setting — 97 tokens per second at max against Sol's 58 — which is why Anthropic's "more than 30% faster output" claim holds. But it takes 12.4 minutes to produce its first token at max effort, against 4.7 minutes for Sol. End to end, Sol finishes an index task in 647 seconds where Opus takes 783.

That has a practical consequence for interactive use. A model's streaming speed is what you feel while reading an answer; its time-to-first-token is what you feel while waiting for one to start. For a chat interface, Opus 5.5's 5.4-second-equivalent settings matter — at medium it starts in 23.5 seconds against Sol's 5.4. For a batch agent that runs unattended, only the total time per task counts, and Sol wins there at every comparable price point.

The same model, three different scores

The most instructive part of this comparison is how much Opus 5.5's numbers move depending on who ran them. These are not errors; they are different harnesses, and they should put a floor under how much weight you give any single figure.

EvaluationAnthropic's own tableOpenAI's launch runIndependent
GDP.pdfnot published28.8% with fallbacks26.2% at max, 28.8% at high
Terminal-Bench-Science 0.158.7%63.3%not in the index
AutomationBench40.0% at max (Zapier)42.5% at max69.5% at max (partial credit)
OSWorld81.8% on 2.0 partial60.3% on 2.0 offlinenot in the index
Intelligence Index v4.3.2not participatednot participated58 at max, 51 at medium

Three things follow. First, AutomationBench is two benchmarks sharing a name: Zapier's binary workflow score and Artificial Analysis's partial-credit version read 40.0% and 69.5% for the same model at the same effort. Quoting either without the qualifier is misleading. Second, OpenAI's GDP.pdf figure for Opus 5.5 excludes nothing but is labelled "with fallbacks", and OpenAI itself notes that its AutomationBench figure for Claude Fable 5.1 omits fallback cost that occurred on roughly 40% of tasks. Third, on Anthropic's side, the Opus 5.5 launch table has no row for GPT-6.1 Sol — it lists GPT-5.6 Sol, the model OpenAI replaced eight days before.

Neither vendor is being dishonest. Both are showing where their model looks best. The reason the Artificial Analysis table anchors this article is simply that it is the only run where both models faced the same tests under the same operator.

Safety and safeguards: two different bets

GPT-6.1 Sol

  • Classified Critical for cybersecurity and High for biological and chemical capability under OpenAI's framework, carrying Astra's safeguard stack.
  • ExploitBench 99.7% at max against GPT-6 Sol's 81.7% — with OpenAI's own caveat that historical-vulnerability contamination may inflate this.
  • Broken-tool honesty improved: when a search or API tool was deliberately degraded, failure to tell the user fell from 4.9% to 2.1% (Astra: 1.5%).
  • Zero bypass attempts in automated safety-reviewer and authorised-scope evaluations.
  • Factual error rate at low effort fell from 11.4% to 7.7% on user-flagged historical errors, staying within 1.9 points of Astra at every level.
  • Instruction following: OpenAI reports clear gains on negative constraints such as "do not modify files outside /src".

Claude Opus 5.5

  • Best score to date on Anthropic's automated behavioural audit — roughly 2,000 scenarios — with containment-boundary attempts about 85% lower than Opus 5, all low severity and self-reported.
  • Transparent rerouting: cyber tasks go to Opus 4.8, biology and frontier-LLM tasks to Opus 5, with the switch visible rather than silent.
  • Gray Swan prompt-injection benchmark: ties Fable 5.1 for the lowest success rate recorded.
  • Vetted access scales through the Cyber Verification Program (three tiers, up to Mythos models) and the Life Sciences Verification Program.
  • EU AI Act watermarking and Zero Data Retention support.
  • However: Anthropic states Opus 5.5 "often suspects it is being evaluated," which it says limits the reliability of its own audit scores.

The fallback policy changes what the benchmarks measure. When Opus 5.5's classifiers fire, work is completed by a different model — and Vals records how often that happened on its index: 91 of 2,281 tasks (3.99%) were fallback-assisted, with 18 refusals (0.79%) scored as failures. On SRE-Bench, outside the index, 82.82% of Opus's tasks were fallback-assisted and 10.30% were refusals, which is why Opus scores 33.59% there against Sol's 50.76%. That is not Opus 5.5 doing security work badly; it is Opus 5.5 declining to do it and handing it to Opus 4.8.

Read positively, that is the safeguard working. Read as a benchmark, it means a "Claude Opus 5.5" row can be a mixture of three models. For security and biology workloads specifically, check your refusal path before comparing scores.

Operational changes to audit before you switch

  • effort: "none" is gone. GPT-6.1 Sol drops the non-reasoning setting that GPT-6 Sol offered. OpenAI's migration guide says to use low instead, so any caller setting none explicitly needs a config change and a regression run.
  • Sol uses 10–30% more output tokens than GPT-6 Sol at each effort setting, per Artificial Analysis. The cost cut came from the rate card, not from efficiency.
  • Sol is slower to stream: 58–59 tokens per second at its top settings against GPT-6 Sol's higher figures and an average of about 70 across tracked models. Latency-sensitive interfaces should be tested before migrating.
  • Opus 5.5 cannot disable thinking at all. If your integration shuts thinking off to control cost, it now returns a 400 — low effort is the replacement.
  • Refusal billing changed on 24 September 2026. Anthropic resumed charging for pre-output refusals in the bio, frontier_llm and reasoning_extraction categories; cyber refusals remain unbilled. If your pipeline triggers safeguards, that is a new line on the invoice — and a fallback is charged separately at the fallback model's rates.
  • Cache minimums differ: 512 tokens for Opus 5.5, 1,024 for Sol. Short system prompts that cache on one model may not cache on the other.
  • Availability: Sol is in the API, ChatGPT Work, Codex, Bedrock and OpenRouter but not in regular ChatGPT chat, and it is off by default in Enterprise and Edu workspaces until an admin enables it. Opus 5.5 is on five cloud platforms from day one.
  • One dead end to name: "cost per accepted task" is the right production metric, and neither vendor publishes it. If Sol attempts a task at $0.40 and passes 40% of the time while Opus attempts at $0.80 and passes 90%, the arithmetic is close to a wash despite a 2× sticker gap. Measure success rates on your own queue.

Verdict

The short version

GPT-6.1 Sol is the better default and Claude Opus 5.5 is the better ceiling, and the line between them is sharp enough to route on. Below index 51, Sol is cheaper at every single level — often by 70–76% for an identical score. Above index 51, Opus 5.5 is the only model that goes there, and the entry price is $1.82 a task, rising to $5.98 at the top.

The decision rule that follows is unusually simple: run Sol by default, and escalate to Opus 5.5 when a task fails a checkpoint — or when you know in advance it needs an index score above 51. Categories that reliably need the ceiling are knowledge-work documents, scientific coding, and open-ended expert reasoning, where Opus's margins are 8 to 13 points on the independent run. Categories that reliably do not are document extraction, high-volume business workflows, and anything short-prompt and output-heavy.

Two caveats before committing. Sol's half-price advantage is real only under 272K input tokens — on large-input, short-output workloads it shrinks to between 1% and 5%. And Sol's own effort ladder peaks at high on coding, so paying for max is likely to buy nothing.

Your workloadRoute toWhy
High-volume extraction, classification, routingGPT-6.1 Sol$0.13 a task at low for the same index score Opus charges $0.55 for
Complex professional PDFs and dense documentsGPT-6.1 SolThe one component it wins outright: 31.0% vs 26.2% on GDP.pdf
Business workflow automationGPT-6.1 Sol at medium35.4% against Opus at 29.5% on OpenAI's run, for about a third of the cost
Budget-capped agent loops with heavy cachingGPT-6.1 Sol$0.10 cached input and the cheapest cost-per-task figure on the frontier
Latency-sensitive interactive agentsGPT-6.1 Sol at lower effort5.4 s to first token at medium against 23.5 s for Opus at the same setting
Terminal and CI coding agentsEither — test bothTerminal-Bench 4.0 reads 59.6% vs 56.1% on one harness and 65.15% vs 55.05% on another
Knowledge work: briefs, decks, analysisClaude Opus 5.5+291 Elo GDPval-AA, +243 Elo AA-Briefcase
Scientific coding and physicsClaude Opus 5.5SciCode 66.9% vs 54.2%; CritPt is a tie at 31.7%
Expert-level reasoning and hard science agentsClaude Opus 5.5 at high+HLE 61.4% vs 52.9%; the only settings that reach index 54–58
Workloads past 272K input on every requestClaude Opus 5.5Flat 1M pricing; Sol's advantage collapses to 1–5%
Authorised offensive securityNeither — different toolsSol is Critical-rated and gated; Opus 5.5 reroutes the work to Opus 4.8
Absolute cost floorGPT-6.1 Sol at low$0.13 a task; nothing cheaper sits at index 42

FAQ

Is GPT-6.1 Sol better than Claude Opus 5.5?

Depends entirely on the effort setting. At low both score 42 and Sol is 76% cheaper. At Opus's default medium (51, $1.34), Sol's xhigh matches the score for $0.39 and Sol's max beats it (52) for $0.72. Above index 51, Opus 5.5 has no competitor: it reaches 54, 56 and 58 where Sol stops at 52. So Sol is better below the crossover and Opus is the only option above it.

Why is Sol exactly half the price?

Input is $2 against $4, output is $10 against $20, and cache writes are $2.50 against $5. Cache reads are the exception where Sol is five times cheaper — $0.10 against $0.20 — because OpenAI cut them when it shipped GPT-6.1 Sol. Sol is not uniformly half price: it is half on most lines, a fifth on cached input, and then more expensive above 272K input because of the long-context surcharge that Opus does not apply.

What is the 272K cliff and does it affect me?

OpenAI applies its long-context rates to the entire request once input exceeds 272,000 tokens: $4 input, $0.20 cache read, $15 output. A request with 272K input and 20K output costs $0.744; at 273K the same request costs $1.392 — an 87% jump for a thousand extra tokens. If your prompts routinely exceed 272K, or your agent appends documents without compaction, assume the half-price advantage does not apply to you. Opus 5.5 serves the full 1M window at flat rates.

Should I use max effort on GPT-6.1 Sol?

Probably not. On OpenAI's DeepSWE v1.1 chart, Sol peaks at high (75.2% for $0.65) and falls to 71.9% at both xhigh and max, the latter costing $1.57 — 2.4× the price for 2.4 points less. On Artificial Analysis's index, the last step from xhigh to max buys one point for 85% more. Ship high for coding and medium for workflows, and test max on your own workload before paying for it.

Does Opus 5.5 cost $4/$20 or $5/$25?

$4 input and $20 output per million tokens, with cached input at $0.20 and cache writes at $5 for five minutes or $8 for an hour. Some third-party price trackers still display $5/$25, which is Claude Opus 5's card — Opus 5.5 cut prices when it shipped on 22 September 2026. Check Anthropic's own pricing page before quoting a number to anyone.

Which model is more honest about its own failures?

They are measured differently and both improved. OpenAI reports that when a search or API tool was deliberately broken, Sol failed to tell the user in 2.1% of runs, down from 4.9% for GPT-6 Sol and close to Astra's 1.5%. On Artificial Analysis's AA-Omniscience, Opus 5.5 scores 46 against Sol's 42 on a scale that rewards correct answers and penalises hallucinations without penalising abstention. Opus 5.5 remains the stronger model on calibration by the independent measure; Sol is the cheaper one.

Can I use both in one pipeline?

Yes, and the ladder makes the rule concrete. Route to Sol by default; escalate to Opus 5.5 when a task fails a checkpoint, or when you know in advance it needs an index score above 51. The two do not share prompt format, tool schemas or reasoning-transcript conventions, so treat the boundary as a full task handoff rather than a mid-conversation model switch — unlike the Opus 5.5 to Fable 5.1 handoff, which does preserve thinking blocks on the same API.

Run the crossover on your own queue

Take fifty tasks you have already solved, run them at Sol's high and Opus 5.5's medium, and count the failures. That single test will tell you whether your work sits above or below the index-51 line.

Open a new chat on CodingFleet →

Sources & further reading

Benchmark scores are vendor- or evaluator-reported and labelled as such throughout. Effort settings are stated wherever the source states them; where a vendor omits the effort level of a comparison, this article says so rather than assuming one. Cross-harness numbers are directional. Prices are USD per million tokens as of early October 2026.