OpenAI replaced GPT-6 Sol with GPT-6.1 Sol seven days after shipping it. That is the shortest gap between a model and its own replacement in this family's history, and it makes for an unusually clean comparison: same price, same context window, same output ceiling, same six-evaluation harness run by the same independent lab. Almost everything that changed can be measured.
What changed is four index points, a halved cache rate, one reasoning setting that no longer exists, and a measurable slowdown. Artificial Analysis, which ran both models across every effort level in one harness, opened its analysis with the line that frames the whole release: "GPT-6.1 Sol replaces GPT-6 Sol after just 7 days."
TL;DR
This is a capability release at the same list price, and it is the first Sol that beats the model it replaces on coding. On Artificial Analysis Intelligence Index v4.3.2, GPT-6.1 Sol gains 4 points at the ceiling (48 → 52) and 7–8 points at every cheaper setting, while cost per task falls at every rung. The gains are largest where you spend the least: low and medium each move 8 points.
- Coding is the headline. Terminal-Bench 4.0 jumps 12 points, the Coding Agent Index gains 3, and on OpenAI's own DeepSWE v1.1 chart the model scores 75.2% at
high— above GPT-6 Astra's best run, for $0.65 a task against Astra's $4.43. - The cache rate halved. Cached input falls from $0.20 to $0.10 per million tokens — a 95% discount on standard input, and half of Claude Sonnet 5.5's rate.
- Three costs. Output tokens rise 10–30%, throughput drops from about 93 to about 67 tokens per second, and time-to-first-token at
maxmore than doubles, from 123 seconds to 268. - One migration. The
nonereasoning setting is gone, and tool calling now requires the Responses API. Anyone using Sol as a cheap non-reasoning function caller has code to change. - GPT-6 Sol is not deprecated. OpenAI lists no retirement for
gpt-6-sol. You are not being forced to move.
+4 in seven days
Terminal-Bench 4.0
$1.06 → $0.72
95% off standard input
Seven days, two Sols
GPT-6 Sol landed on 22 September 2026 as a price cut wearing a version number: it matched GPT-5.6 Sol's intelligence almost exactly while halving the price. The reception was mixed, and the criticism was specific — on OpenAI's own charts, GPT-6 Sol's best DeepSWE and OSWorld scores sat below GPT-5.6 Sol's, which is an awkward thing for a replacement to do. A Hacker News commenter put it plainly at the time: "6 Sol performs worse than 5.6 Sol at DeepSWE? Weird!!"
GPT-6.1 Sol, announced at DevDay on 29 September 2026, fixes that. It clears both of last week's regressions, and it does it without moving the price. OpenAI's framing is "near-Astra intelligence for a fifth of the price," and the independent numbers support a version of that claim.
| Specification | GPT-6.1 Sol | GPT-6 Sol | Change |
|---|---|---|---|
| Released | 29 September 2026 DevDay | 22 September 2026 | +7 days |
| API model ID | gpt-6.1-sol | gpt-6-sol | — |
| Input | $2.00 / MTok | $2.00 / MTok | unchanged |
| Cached input | $0.10 / MTok | $0.20 / MTok | −50% |
| Cache write | $2.50 / MTok | $2.50 / MTok | unchanged |
| Output | $10.00 / MTok | $10.00 / MTok | unchanged |
| Above 272K input, in / cache / out | $4 / $0.20 / $15 | $4 / $0.20 / $15 | unchanged |
| Batch and Flex | 50% off — $1 / $5 | 50% off — $1 / $5 | unchanged |
| Context window | 1,050,000 tokens | 1,050,000 tokens | unchanged |
| Max output | 128,000 tokens | 128,000 tokens | unchanged |
| Knowledge cutoff | 30 April 2026 | 20 April 2026 | +10 days |
| Effort levels | low · medium · high · xhigh · max | none · low · medium · high · xhigh · max | none removed |
| Default effort | medium | medium | unchanged |
| Tool calling | Responses API required | Responses API, or Chat Completions with effort: none | narrower |
| Where it runs | ChatGPT Work, Codex, OpenAI API, Amazon Bedrock | same | unchanged |
| Ultrafast tier | announced, "coming soon" — up to 8× in Codex | not offered | new |
| System card | addendum published | addendum published | unchanged |
| Deprecation status | current | still available, no retirement announced | — |
The cache line is the only pricing change, and it is the one that matters most for agents. At $0.10 per million tokens, a cached read costs 5% of a fresh read — a 95% discount. For a Codex-style loop that re-reads the same repository context on every step, that is the single largest lever on the bill, and it compounds with last week's steerable caching, which lets you change reasoning effort or tool availability mid-conversation without invalidating the cache.
The ladder: eleven settings, one harness
This is the table the comparison should be built on. Both models ran the identical ten evaluations under Artificial Analysis — AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR v1.1 — and cost is measured per task including cache reads and writes. GPT-6 Sol's none setting is included because it is the row that no longer exists.
| Configuration | Index | Cost / task | Score delta | Cost delta |
|---|---|---|---|---|
GPT-6 Sol · nonereasoning switched off — removed in 6.1 | 28 | $0.33 | no equivalent | — |
GPT-6.1 Sol · low | 42 | $0.13 | +8 | 0% |
GPT-6 Sol · low | 34 | $0.13 | — | — |
GPT-6.1 Sol · mediumboth models' default | 48 | $0.21 | +8 | −16% |
GPT-6 Sol · medium | 40 | $0.25 | — | — |
GPT-6.1 Sol · high | 50 | $0.32 | +7 | −14% |
GPT-6 Sol · high | 43 | $0.37 | — | — |
GPT-6.1 Sol · xhigh | 51 | $0.39 | +7 | −26% |
GPT-6 Sol · xhigh | 44 | $0.53 | — | — |
GPT-6.1 Sol · max | 52 | $0.72 | +4 | −32% |
GPT-6 Sol · max | 48 | $1.06 | — | — |
Read the deltas in order and a pattern appears that the headline hides. The gain is largest at the cheap end and smallest at the top: +8 at low and medium, +7 at high and xhigh, and only +4 at max. Meanwhile the cost saving runs the other way — nothing at low, −16% at medium, −32% at max. The two effects are complementary: the model improved most where you spend least, and got cheapest where you spend most.
medium — its default — scores 48 for $0.21. GPT-6 Sol needed $1.06 to reach 48. That is a five-fold cost reduction for the same index score. Source: Artificial Analysis Intelligence Index v4.3.2.The row that no longer exists
GPT-6 Sol's none setting scored 28 for $0.33 a task. GPT-6.1 Sol has no equivalent, and OpenAI's migration guide says to use low instead — which scores 42 for $0.13. So the replacement is 14 index points better and 61% cheaper than the setting it removed. That is a rare case where a breaking change is also an upgrade, and it is worth saying plainly because the migration itself is real work.
Where the gains are, component by component
The composite hides a lot, so here is every component Artificial Analysis publishes for GPT-6.1 Sol, at all five effort levels. Two of these columns are non-monotonic, which matters for how you configure the model.
| Evaluation | low | medium | high | xhigh | max |
|---|---|---|---|---|---|
| AA-Briefcase v1.1agentic knowledge work · Elo | 1,118 | 1,365 | 1,471 | 1,507 | 1,564 |
| GDPval-AA v2.1real-world work tasks · Elo | 1,297 | 1,433 | 1,486 | 1,510 | 1,575 |
| AutomationBench-AASaaS workflows · partial credit | 52.6% | 62.6% | 64.5% | 66.6% | 64.9% |
| Terminal-Bench 4.0agentic coding & terminal use | 30.8% | 48.0% | 51.5% | 54.0% | 56.1% |
| SciCodescientific coding | 53.2% | 53.2% | 55.8% | 55.7% | 54.2% |
| Legal Index | 50 | 54 | 56 | 57 | 58 |
| Engineering Index | 44 | 50 | 52 | 53 | 54 |
| Economics Index | 52 | 56 | 58 | 58 | 59 |
Three of these peak below max. AutomationBench-AA is highest at xhigh (66.6%) and falls to 64.9% at max. SciCode is highest at high (55.8%) and drops to 54.2% at max. The Economics Index is flat between high and xhigh before gaining a single point at max. On those three evaluations, turning the dial further costs more and returns less.
The two Elo benchmarks behave the opposite way and climb all the way to max. So there is no single correct setting: the effort level that maximises a coding score is not the one that maximises a business-workflow score, and on this model the gap is large enough to matter.
What moved against GPT-6 Sol
Artificial Analysis published the deltas directly. These are index points unless stated otherwise, measured at matched effort.
| Evaluation | Change vs GPT-6 Sol | Note |
|---|---|---|
| Intelligence Index | +4 | 48 → 52 at max; +7 to +8 at every cheaper setting |
| Terminal-Bench 4.0 | +12 | The single largest gain in the release |
| AA-Omniscience accuracy | +8 | With hallucination rate falling from 60% to 54% |
| GDP.pdf | +6 | Professional document reasoning over dense PDFs |
| Humanity's Last Exam | +5 | Expert-level multidisciplinary reasoning |
| GDPval-AA v2.1 | +5 | Real-world work tasks across occupations |
| AA-Briefcase v1.1 | +4 | About +80 Elo, driven by rubric score and analytical quality |
| Coding Agent Index | +3 | At max; 2 points below GPT-6 Astra |
| Output tokens per task | +10% to +30% | The model works harder for its score |
One regression inside the gain. Artificial Analysis notes that GPT-6.1 Sol's AA-Briefcase improvement comes from "increases in its rubric score and Analytical Quality Elo, while Presentation Elo falls slightly." That is the same failure mode that hurt GPT-6 Sol's knowledge-work numbers — shorter deliverables that skip required components — partially fixed but not fully. If your output is a client document with a required structure, test it rather than assuming the +80 Elo covers you.
Coding: the DeepSWE chart, read properly
OpenAI's DeepSWE v1.1 chart is the most useful vendor data in the release, because it publishes absolute scores and cost per task at all five effort levels for three models. Here it is in full.
| Effort | GPT-6.1 Sol | GPT-6 Sol | GPT-6 Astra |
|---|---|---|---|
| low | 64.4% · $0.17 | 37.2% · $0.16 | 67.0% · $1.60 |
| medium | 73.0% · $0.42 | 56.6% · $0.38 | 72.8% · $3.08 |
| highthe peak for GPT-6.1 Sol | 75.2% · $0.65 | 65.3% · $0.64 | 73.2% · $3.92 |
| xhigh | 71.9% · $0.79 | 66.6% · $1.00 | 74.1% · $4.43 |
| max | 71.9% · $1.57 | 68.8% · $2.74 | 73.2% · $7.50 |
Three things fall out of this table. First, GPT-6.1 Sol's best score of 75.2% beats GPT-6 Astra's best run of 74.1% — and OpenAI's own summary is that it does so at "approximately 76% lower cost per task." Second, the biggest single improvement over GPT-6 Sol is at low effort, where the score nearly doubles from 37.2% to 64.4% for one cent more per task. Third, and least comfortable for OpenAI, the score declines after high: 75.2% at high, then 71.9% at both xhigh and max, where the task cost more than doubles to $1.57.
high and then declines. The two faded bars are the settings that cost more and score less. GPT-6 Sol's curve climbs monotonically, so a team that tuned its effort level on the old model will be over-spending on the new one. Source: OpenAI DeepSWE v1.1 chart data.The regression nobody put in the headline: speed
Every other change in this release is an improvement or a neutral. Speed is not, and it is the one number OpenAI's launch material does not lead with.
| Measure | GPT-6.1 Sol | GPT-6 Sol | Direction |
|---|---|---|---|
Output speed at max | ≈67 tokens/s | ≈93 tokens/s | ≈28% slower |
Output speed at low | 74 tokens/s | ≈104 tokens/s | slower |
Time to first token at max | 267.6 s | 123.1 s | 2.2× longer |
Time to first token at low | 1.84 s | not published | fast |
| Output tokens per task | +10% to +30% | baseline | more |
There is a mitigation, and it is the same one the coding data points to: run at a lower effort level. GPT-6.1 Sol at low starts answering in 1.84 seconds and streams at 74 tokens per second — faster than either model at max — while still scoring 42 on the index, which is 8 points above GPT-6 Sol at low and 14 above GPT-6 Sol's removed none setting.
The migration: one setting, one endpoint
Two changes will break code that runs against GPT-6 Sol today. Both are documented, and neither is hard, but both need a regression test.
| Change | What breaks | Fix |
|---|---|---|
The none reasoning setting is gone | Any request setting reasoning.effort to none — or minimal, which is also unsupported — fails. GPT-6 Sol accepted none; GPT-6.1 Sol supports only low, medium, high, xhigh and max. | OpenAI's migration guide says to use low. It is 14 index points better and 61% cheaper than the setting it replaces, so this is an upgrade disguised as a breaking change. |
| Tool calling requires the Responses API | GPT-6 Sol allowed function calling from Chat Completions as long as reasoning was set to none. GPT-6.1 Sol has no none, so that path is closed. Plain Chat Completions text calls still work. | Move tool-calling workloads to the Responses API. This is the change that costs real engineering time. |
| Reasoning is always on | Nothing fails, but a model that always thinks is a model that always bills reasoning tokens. Budgets tuned on GPT-6 Sol's none path will not hold. | Re-measure cost per task at your chosen effort level rather than carrying a budget over. |
The one that will surprise people. If you used Sol as a cheap non-reasoning function caller — a classifier, a router, a structured extractor — that pattern is gone. The replacement is low effort on the Responses API, which is better and cheaper per task but is a different request shape. Budget an afternoon, not a string swap.
What the halved cache rate is worth
Because input, output and cache-write rates are unchanged, the only pricing difference is the cache read line — and its value depends entirely on how much of your input is cached. Here is the illustrative cache-heavy agent task used throughout this series, priced at both cards.
| Line, per million tokens | GPT-6.1 Sol | GPT-6 Sol | Change |
|---|---|---|---|
| Cache reads · 8,000,000 | $0.80 | $1.60 | −50% |
| Uncached input · 400,000 | $0.80 | $0.80 | — |
| Cache writes · 600,000 | $1.50 | $1.50 | — |
| Output · 300,000 | $3.00 | $3.00 | — |
| Total | $6.10 | $6.90 | −12% |
At equal token counts the cache cut is worth 12% on a cache-heavy task — smaller than the headline suggests, because cache reads are only one of four lines. The measured figure is larger, a 32% reduction at max, and the difference is efficiency: GPT-6.1 Sol reaches a higher score in fewer effective steps, so it spends less on the other three lines too. Real workloads also use 10–30% more output tokens, which pushes back the other way.
Where the cache cut pays most is at the extreme: a pipeline whose input is 95% cached reads sees almost the full 50% reduction on its dominant cost line. OpenAI's own worked example from the GPT-6 Sol launch — a codebase migration agent — ran end to end for $0.7082 with 91% of input tokens served from cache. On GPT-6.1 Sol's card the same run's cache line halves.
What OpenAI claims, and what the independent run says
Both sets of numbers are in this article, and they agree on direction while differing on precision. Here they are side by side so you can see which is which.
| Claim | OpenAI's figure | Independent figure |
|---|---|---|
| Coding | DeepSWE v1.1 75.2% at high, above Astra's best 74.1%, at ~76% lower cost per task | Terminal-Bench 4.0 +12 points; Coding Agent Index +3 at max, 2 below Astra |
| Computer use | OSWorld 2.0 offline 71.4% at max, +7.0 over GPT-6 Sol, −2.1 under Astra at ~1/7 the cost | Not in the Intelligence Index; no independent OSWorld run published |
| Professional documents | GDP.pdf 32.0% at high, +4.0 over GPT-6 Sol, −0.2 under Astra | GDP.pdf +6 points vs GPT-6 Sol |
| Business workflows | AutomationBench 1.0.6 36.1% at max, +2.9 over GPT-6 Sol; at medium, +4.8 and +2.2 over Opus 5.5 | AutomationBench-AA 66.6% at xhigh, above its own max of 64.9% |
| Science | Terminal-Bench Science 0.1 57.0% at max, more than double GPT-6 Sol, at $5.47 a task against Astra's $23.80 | Not in the Intelligence Index |
| Factuality | Error rate at low falls from 11.4% to 7.7%; within 1.9 points of Astra at every setting | AA-Omniscience accuracy +8 points; hallucination 60% → 54% |
| Safety | No attempts to bypass the automated safety reviewer; fails to flag a broken search tool in 2.1% of cases vs 4.9% for GPT-6 Sol | Not independently measured |
One vendor claim deserves a caveat. OpenAI's AutomationBench win over Claude Opus 5.5 is a medium-versus-medium comparison: 31.7% against 29.5%, a 2.2-point lead. At max effort the ordering reverses — Opus 5.5 scores 42.5% and Astra 41.4% against GPT-6.1 Sol's 36.1%. The claim is accurate as stated and misleading if you read it as a general result. On the same chart, GPT-6.1 Sol's max score costs $0.30 a task against Opus 5.5's $1.44, so the value argument holds even where the score does not.
What people said in the first day
The Hacker News launch thread passed 1,000 points and 900 comments in a day, and the split was roughly "big step up from 6 Sol" against "still not Opus."
"Sol 6.1 is very noticeably smarter than sol 6 even after half a day of using it."
— baq, Hacker News
"It's the same for most tasks. Where I do notice it is long agent runs, where agents take more steps and the performance difference definitely compounds over the iterations."
— kdaniel_03, Hacker News
The second quote is the more useful one, and it matches the benchmark shape: the gains are largest at the cheap end of the ladder, which is exactly where long agent loops spend their steps. A commenter running a cost comparison against DeepSeek V4.1 Flash reported GPT-6.1 Sol at medium finishing in 2.2 minutes for $0.21 and 8,000 tokens, against DeepSeek's 5.5 minutes for $0.27 and 89,000 tokens — a token-efficiency gap of more than tenfold on that task.
Two criticisms recurred. The first is that GPT-6 Sol should never have shipped: "6 Sol was worse than 5.6 Sol from my own experiences. Far worse. Will see if this remedies things." The second is not about the model at all — OpenAI cut Codex subscription allowances the same week, with Pro 200 dropping from 20× to 10× Plus usage, so for subscribers a cheaper model does not translate into more work per dollar.
Verdict
The short version
Migrate, and expect the migration to be worth an afternoon. GPT-6.1 Sol is a genuine capability gain at an unchanged list price, and unlike last week's release it beats the model it replaces on the benchmarks that were embarrassing — DeepSWE and OSWorld both clear GPT-5.6 Sol's best runs now. The ceiling moves 48 → 52, the cheap settings move 7–8 points, and cost per task falls at every rung.
The costs are real and specific: throughput drops about 28%, time-to-first-token at max more than doubles, output tokens rise 10–30%, the none setting is gone, and tool calling now requires the Responses API. If you run an interactive product, measure the latency before you switch. If you run a batch agent, the arithmetic is straightforward.
The one thing you should not do is carry your old effort setting over. GPT-6 Sol's curve climbed monotonically; GPT-6.1 Sol's peaks at high on coding and xhigh on business workflows, and declines after. Re-sweep effort or you will pay more for a lower score.
| Your workload | Route to | Why |
|---|---|---|
| Coding agents, Codex, complex refactors | GPT-6.1 Sol at high | 75.2% on DeepSWE, above Astra's best, at $0.65 a task; the peak of the curve |
| Business workflow automation | GPT-6.1 Sol at xhigh | AutomationBench-AA peaks at 66.6%, above its own max of 64.9% |
| High-volume tagging, routing, extraction | GPT-6.1 Sol at low | 42 on the index for $0.13, 1.84s to first token, 74 tokens/s |
| Long agent loops that re-read context | GPT-6.1 Sol | $0.10 cache reads, half the old rate, with steerable caching |
| Document-heavy professional work | GPT-6.1 Sol at high | GDP.pdf +6 points independently, 32.0% on OpenAI's run |
| Interactive chat where latency matters | Test first | 267.6s to first token at max, against 123.1s for GPT-6 Sol |
| Cheap non-reasoning function calling | Re-engineer | none is gone and tools need the Responses API |
| Peak score on hard science | GPT-6 Astra | 68.1% on Terminal-Bench Science 0.1 against Sol's 57.0% |
| Nothing at all | GPT-6 Sol | Not deprecated; if your workload passes, there is no forced move |
FAQ
Is GPT-6.1 Sol actually better than GPT-6 Sol?
On every evaluation Artificial Analysis measured, yes. The Intelligence Index ceiling moves 48 → 52, and the gains are larger at cheaper settings: +8 at low and medium, +7 at high and xhigh, +4 at max. Terminal-Bench 4.0 gains 12 points, AA-Omniscience accuracy gains 8 with hallucination falling from 60% to 54%, and the Coding Agent Index gains 3. The exceptions are speed and token use, both of which got worse.
Why did OpenAI replace a model after seven days?
Because GPT-6 Sol regressed on the benchmarks OpenAI had been using to sell it. On OpenAI's own charts, GPT-6 Sol's best DeepSWE score of 68.8% sat below GPT-5.6 Sol's, and its OSWorld score did too. GPT-6.1 Sol clears both: 75.2% on DeepSWE at high, above GPT-6 Astra's best run, and 71.4% on OSWorld 2.0 offline. Artificial Analysis's summary of the replacement is blunt: "GPT-6.1 Sol replaces GPT-6 Sol after just 7 days."
What reasoning effort should I use?
Not max, on most workloads. Start at high for coding — that is where DeepSWE peaks at 75.2% before declining to 71.9% at xhigh and max — and xhigh for business workflows, where AutomationBench-AA peaks at 66.6% against 64.9% at max. Use low for high-volume work: 42 on the index for $0.13, with 1.84-second time-to-first-token. The one setting to test rather than assume is max.
Is the halved cache rate a big deal?
It depends entirely on your cache-hit ratio. On a task where 8M of 9M input tokens are cache reads, the cut is worth 12% of the total bill at equal token counts. On a pipeline that is 95% cached reads, it is worth close to 50% on the dominant line. OpenAI's own worked codebase-migration example ran at 91% cached input, so for agent loops that re-read a repository or document set it is the largest single lever in the release.
Do I have to migrate?
No. OpenAI's deprecations page lists no retirement for gpt-6-sol — the models being shut down in December 2026 are the original GPT-5 snapshots and o3, which point at gpt-5.6-sol. GPT-6 Sol keeps working at its current price. The argument for moving is that GPT-6.1 Sol is better at every effort level and cheaper per task at every effort level where costs differ, so staying is a choice to pay more for less.
Why is GPT-6.1 Sol slower than the model it replaces?
OpenAI has not explained it, and the data is unambiguous: about 67 tokens per second at max against about 93 for GPT-6 Sol, with time-to-first-token at max rising from 123 seconds to 268. It also uses 10–30% more output tokens. The likely mechanism is more reasoning per step, which would also explain the capability gain — but that is inference, not something either company has stated. Artificial Analysis rates the model "slower than average" among those it tracks.
Re-sweep your effort level before you switch
The single most expensive mistake available here is carrying a GPT-6 Sol effort setting onto GPT-6.1 Sol. The old curve climbed to max; the new one peaks at high and declines. Run your own tasks at three settings and pick the one that wins.
Sources & further reading
- Artificial Analysis — GPT-6.1 Sol replaces GPT-6 Sol after just 7 days: the four-point index gain, component deltas, Coding Agent Index, token efficiency and AA-Omniscience findings.
- Artificial Analysis — GPT-6.1 Sol release page and the GPT-6 Sol release page: effort ladders, cost per task, output speed.
- Artificial Analysis — GPT-6.1 Sol (max) vs GPT-6 Sol (max): component scores, speed and latency side by side.
- OpenAI — Introducing GPT-6.1 Sol: DeepSWE, OSWorld, GDP.pdf, AutomationBench, Terminal-Bench Science and factuality claims.
- OpenAI — API deprecations: retirement schedule and replacement mappings.
- Emergent — GPT-6.1 Sol benchmarks: the DeepSWE effort table, the
noneremoval, and the speed regression. - Handy AI — Model Drop: GPT-6.1 Sol: per-benchmark deltas, the medium-versus-max AutomationBench caveat, and the Hacker News reaction round-up.
- eesel — GPT-6.1 Sol review: 75 support-ticket tests against Astra, and the Responses API migration in practice.
- DataCamp — GPT-6.1 Sol: features, benchmarks, pricing, access: the supported effort levels and the tool-calling requirement.
- Related on CodingFleet: GPT-6 Sol vs GPT-5.6 Sol, Claude Opus 5.5 vs GPT-6.1 Sol, Claude Opus 5.5 vs GPT-6 Sol, GPT-6 Astra Review.
Benchmark scores are vendor- or evaluator-reported and labelled as such throughout. Effort settings are stated wherever the source states them; where a vendor omits the effort level of a comparison, the article says so rather than assuming one. Cross-harness numbers are directional. Prices are USD per million tokens as of early October 2026.