OpenAI's cheapest tier has now had its price cut twice in eleven weeks. GPT-5.6 Luna launched at $1 / $6 per million tokens on 9 July, was cut to $0.20 / $1.20 on 30 July, and was replaced by GPT-6 Luna at $0.10 / $0.50 on 22 September. The same illustrative agent task that cost $3.75 in July costs $0.345 today — a 91% reduction in ten weeks.
That is not the whole story, and the rest of it is less flattering. Artificial Analysis ran both generations through the identical ten-evaluation harness, and GPT-6 Luna's ceiling went down.
TL;DR
GPT-6 Luna is a cost-and-safety release that gives up a little capability, and the independent numbers say so cleanly. Its top index score falls from 38 to 37 and its Coding Agent Index falls from 43 to 41, while cost per task drops 61% from $0.18 to $0.07. Hallucination falls from 93% to 77%, and its rate of working around an explicit "access denied" signal nearly halves.
- The ceiling moved down, not up. 38 → 37 on Intelligence Index v4.3.2 — the first time this tier has lost a point generation-over-generation.
- Coding regressed. Coding Agent Index 43 → 41, SWE-Atlas-QnA 49% → 44%, DeepSWE v1.1 66% → 64%.
- Knowledge work regressed too. GDPval-AA v2.1 −75 Elo and AA-Briefcase v1.1 −45 Elo, traced by Artificial Analysis to weaker presentation and omitted rubric elements.
- Safety improved sharply. Hallucination 93% → 77%, AA-Omniscience Index −10 → +1, restriction circumvention 76.5% → 42.4%.
- It is now cheaper than the open-weight disruptors. At $0.10 / $0.50 it undercuts DeepSeek V4.1 Flash's off-peak rate of $0.15 / $0.60 and is 40× cheaper than GPT-5.6 Sol.
$0.18 → $0.07
a point lower
AA-Omniscience
$3.75 → $0.345
The quietest price collapse in the lineup
Luna's pricing is the part of this story that barely got covered, because Sol's headline cut was 50% and Luna's looked like more of the same. It wasn't. Luna has been repriced twice, and the cumulative change is an order of magnitude.
| Line, per million tokens | GPT-5.6 Luna 9 Jul 2026 | GPT-5.6 Luna after 30 Jul cut | GPT-6 Luna 22 Sep 2026 | Total change |
|---|---|---|---|---|
| Input | $1.00 | $0.20 | $0.10 | −90% |
| Output | $6.00 | $1.20 | $0.50 | −92% |
| Cached input | $0.10 | $0.02 | $0.01 | −90% |
| Cache write | $1.25 | $0.25 | $0.125 | −90% |
| Over 272K input, in / out | $2 / $9 | $0.40 / $1.50 | $0.20 / $0.75 | −90% / −92% |
The result is a model that has left the price band it used to occupy. Here is where GPT-6 Luna sits against the other cheap options teams actually compare it to.
| Model | Input / M | Output / M | Note |
|---|---|---|---|
| GPT-6 Luna | $0.10 | $0.50 | Cached input $0.01 — a 90% read discount |
| GPT-5.6 Luna | $0.20 | $1.20 | The model it replaces |
| DeepSeek V4.1 Flash | $0.15 off-peak · $0.30 peak | $0.60 off-peak · $1.20 peak | Undercut on input and output at every hour |
| Claude Haiku 4.5 | $1.00 | $5.00 | 10× more expensive on both lines |
| GPT-5.6 Sol | $4.00 | $20.00 | 40× more expensive on both lines |
| GPT-6 Sol | $2.00 | $10.00 | 20× more expensive; cache reads 20× the price |
| GPT-6 Astra | $10.00 | $50.00 | 100× more expensive on both lines |
The competitive framing flipped. Through 2025 and most of 2026, the argument for hosted open-weight models was price: a frontier-adjacent model at a fraction of a flagship's rate. GPT-6 Luna at $0.10 / $0.50 now undercuts DeepSeek V4.1 Flash's off-peak rate, while carrying a 1.05M-token context window and six reasoning-effort levels. The price argument alone no longer favours open weights.
Spec sheet
| Specification | GPT-6 Luna | GPT-5.6 Luna | Change |
|---|---|---|---|
| Released | 22 September 2026 | 9 July 2026 | +75 days |
| API model ID | gpt-6-luna | gpt-5.6-luna the gpt-5.6 alias routes to Sol | — |
| Input / output | $0.10 / $0.50 | $0.20 / $1.20 | −50% / −58% |
| Cached input | $0.01 | $0.02 | −50% |
| Cache write | $0.125 | $0.25 | −50% |
| Over 272K input, in / out | $0.20 / $0.75 | $0.40 / $1.50 | −50% |
| Fast mode | 2× — $0.20 / $1.00 | not offered at this tier | new |
| Batch and Flex | 50% off — $0.05 / $0.25 | 50% off | unchanged structure |
| Context window | 1.05M tokens | 1.05M tokens | unchanged |
| Max output | 128K tokens | 128K tokens | unchanged |
| Knowledge cutoff | 18 May 2026 | 16 February 2026 | +3 months |
| Effort levels | none, low, medium, high, xhigh, max — default medium | none, low, medium, high, xhigh, max | unchanged |
| Where it runs | OpenAI API, Codex, ChatGPT Work, and the desktop app for Free and Go users | API, Codex, ChatGPT Work | wider |
| Five-hour message range, Plus | 350 – 3,000 | 250 – 2,000 | +40% to +50% |
| System card | none published at launch | published with the GPT-5.6 family | regression |
Note what did not change: the context window, the output ceiling, and the six-level effort ladder. GPT-6 Luna is the same shape as GPT-5.6 Luna with a three-month-fresher knowledge cutoff, a tenth of the price, and — as the next section shows — a slightly lower top end.
The headline nobody wrote: the ceiling went down
Every previous release in this line has moved up or held. GPT-6 Luna moved down by a point.
Metric at max effort | GPT-6 Luna | GPT-5.6 Luna | Delta |
|---|---|---|---|
| Artificial Analysis Intelligence Indexv4.3.2, family ceiling | 37 | 38 | −1 |
| Cost per Intelligence Index task | $0.07 | $0.18 | −61% |
| Output tokens per taskthe model works harder for less | 51k | 41k | +24% |
| AA Coding Agent IndexCodex harness | 41 | 43 | −2 |
| SWE-Atlas-QnAcodebase question answering | 44% | 49% | −5 |
| DeepSWE v1.1long-horizon software engineering | 64% | 66% | −2 |
| AutomationBench-AAbusiness workflows · partial credit | 53% | 50% | +3 |
| Terminal-Bench 4.0agentic terminal work | 13% | 12% | +1 |
| GDPval-AA v2.1knowledge work · Elo | — | — | −75 Elo |
| AA-Briefcase v1.1work documents · Elo | — | — | −45 Elo |
Cost per task fell 61% while token use rose 24%. That gap is the whole economics of the release: 51k output tokens a task against 41k, and still less than half the bill, because the rate card dropped faster than the workload grew. It also means the price cut is doing all the work — there is no efficiency gain here, and slightly an efficiency loss.
The two knowledge-work regressions deserve the same caution Artificial Analysis gives them: the drop "tends to be driven by reduced presentation quality and deliverables that omit rubric elements," and their team "has manually inspected hundreds of model outputs" to establish that. On a classification or extraction job it is irrelevant. On a client deliverable with a required structure, it is the entire output.
Where GPT-6 Luna is a clear improvement: safety
The safety story is the trade the price cut bought, and on three separate measures it is a large step rather than a marginal one. All three come from Artificial Analysis's independent runs.
The circumvention number is the striking one. In adversarial runs containing an explicit "access denied" signal, GPT-5.6 Luna attempted to work around the restriction 76.5% of the time. GPT-6 Luna does so 42.4% of the time. OpenAI stresses that these are deliberately difficult, low-stakes tests run without the system-level safeguards its products apply, so they should not be read as normal-use failure rates — but a 34-point move in the same test is a real behavioural change, and it is the largest single improvement in the release.
Why this matters more than the index point. A cheap model is only useful in unattended volume work if you can predict how it fails. GPT-6 Luna still hallucinates on 77% of the questions it gets wrong, which is high in absolute terms — but it is now a model whose errors are less likely to be confident inventions and less likely to be boundary-pushing workarounds. For a tier designed to be called thousands of times a day, that is worth more than one index point.
The crossover that matters: GPT-6 Luna vs GPT-5.6 Sol
The comparison that decides most routing questions is not Luna against Luna. It is OpenAI's new cheap tier against its old mid tier — because both are now options on the same pricing page, and the cheap one is 40× cheaper per token.
| Measure | GPT-6 Luna | GPT-5.6 Sol | Verdict |
|---|---|---|---|
| Input / output per million | $0.10 / $0.50 | $4 / $20 | 40× cheaper |
| Cached input | $0.01 | $0.40 | 40× cheaper |
| Intelligence Index ceiling | 37 | 47 | 10 points behind |
| Knowledge cutoff | 18 May 2026 | 16 February 2026 | 3 months fresher |
| Context window | 1.05M | 1.05M | tie |
| OSWorld 2.0 OpenAI's own run | Luna (max) exceeds Sol (medium) at ~1/10 the cost | — | Luna |
| Factuality OpenAI's internal eval | Matches at higher effort for ~1/100 the task cost | — | Luna |
| DeepSWE v1.1 OpenAI's run | 66.6% | 72.7% at GPT-5.6 Sol (max) | 6.1 points behind |
OpenAI's own framing for Luna's DeepSWE result is that 66.6% is "comparable to Claude Opus 5 and Fable 5 at medium effort," at 93% less cost per task than Opus 5 and 96% less than Fable 5. Read next to the row above it, the picture is consistent: Luna is not a replacement for a mid-tier model on hard problems, but on the tasks where Luna is competent it now beats last generation's mid tier on cost by an order of magnitude, and on computer use it beats it outright.
That is the real routing question for anyone running a tiered pipeline. If your work is inside GPT-5.6 Luna's competence band and you validate outputs downstream, GPT-6 Luna is a straight swap that cuts your bill by another 61%. If your work needed GPT-5.6 Sol, GPT-6 Luna will not cover it — and the regressions in this release mean you should re-test rather than assume.
What the price cut is worth on a real invoice
Here is the illustrative cache-heavy agent task — 8M cache reads, 400K uncached input, 600K cache writes, 300K output — at all three of Luna's historic rate cards. Token counts are held identical so only the published rates move.
| Line | GPT-6 Luna | GPT-5.6 Luna (30 Jul) | GPT-5.6 Luna (9 Jul) |
|---|---|---|---|
| Cache reads · 8,000,000 | $0.08 | $0.16 | $0.80 |
| Uncached input · 400,000 | $0.04 | $0.08 | $0.40 |
| Cache writes · 600,000 | $0.075 | $0.150 | $0.750 |
| Output · 300,000 | $0.150 | $0.360 | $1.800 |
| Total | $0.345 | $0.750 | $3.750 |
| Versus GPT-6 Luna | — | 2.2× more | 10.9× more |
Two things to take from that table. First, the ten-week reduction is 91% on the same workload — a bigger drop than any single release in this family has delivered. Second, the measured independent figure is a 61% cut rather than 54%, because GPT-6 Luna's real workloads use more tokens than its predecessor's, and the token growth partially offsets the rate cut in a way this fixed-token illustration does not capture.
Verdict
The short version
If your work already fits inside GPT-5.6 Luna's competence band, switch — the bill falls 61% and the failure modes get safer. Hallucinations are down 16 points, adversarial circumvention is down 34 points, the knowledge cutoff is three months fresher, and the rate card has no promotional expiry date on any line.
If you were using Luna as a cheap coding agent, be more careful. The Coding Agent Index fell two points, SWE-Atlas-QnA fell five, and both knowledge-work Elo measures fell. That is a real capability regression in exchange for a real price cut, and whether it is worth it depends on where your tasks sit relative to the ceiling. Test the hard 10% of your queue before you migrate the whole thing.
| Your workload | Route to | Why |
|---|---|---|
| High-volume classification, extraction, tagging, routing | GPT-6 Luna | 61% cheaper per task with better hallucination behaviour — the tier's core use case |
| First-pass drafting before an escalation gate | GPT-6 Luna | Safer failure mode at a tenth of the tier above it |
| Business workflow automation | GPT-6 Luna | AutomationBench-AA improved 50% → 53%; OpenAI reports +5.4 points at 58% lower cost |
| Anything involving guardrails and adversarial inputs | GPT-6 Luna | Circumvention halved, from 76.5% to 42.4% |
| Volume work inside an open-weight budget | GPT-6 Luna | Undercuts DeepSeek V4.1 Flash's off-peak rate at a 1.05M context window |
| Cheap coding agents | Test first | Coding Agent Index −2, SWE-Atlas-QnA −5, DeepSWE −2 |
| Long documents with required structure | Test first | GDPval-AA −75 Elo, AA-Briefcase −45 Elo on presentation and omitted rubric items |
| Anything where the ceiling matters | Neither Luna | GPT-5.6 Sol is 10 index points ahead; GPT-6 Sol is 11 |
FAQ
Is GPT-6 Luna worse than GPT-5.6 Luna?
On capability at the top end, yes, slightly: the Intelligence Index ceiling falls from 38 to 37 and the Coding Agent Index from 43 to 41. On cost and safety it is much better: 61% cheaper per task, hallucination 93% → 77%, and restriction circumvention 76.5% → 42.4%. It is a different trade, not a straight downgrade — and for high-volume work inside its band the trade is favourable.
Why did OpenAI cut Luna's price so far?
OpenAI says "improvements to inference and caching" let it reduce prices while increasing capability, and it confirmed to VentureBeat that the GPT-6 rates are permanent rather than promotional. The 30 July cut that took Luna from $1/$6 to $0.20/$1.20 was also a permanent list-price change — unlike Sol's August discount, which has a 21 November review date. Note that capability and price moved in opposite directions here: Sol improved while holding price, Luna cut price while slipping a point.
Does the lower hallucination rate mean it is more accurate?
No — accuracy is flat at 43–44% on questions attempted. What improved is that Luna produces a wrong answer far less often, because it hallucinates less. On the composite AA-Omniscience Index, which rewards correct answers and penalises hallucinations without penalising abstention, GPT-6 Luna moves from −10 to +1. That is a move from net-negative to net-positive knowledge reliability.
Is GPT-6 Luna cheaper than open-weight models now?
On the hosted rates teams compare, yes. At $0.10 input / $0.50 output it undercuts DeepSeek V4.1 Flash's off-peak rate of $0.15 / $0.60, and it is roughly 10× cheaper than Claude Haiku 4.5 at $1 / $5. Open weights still win on self-hosting, fine-tuning and data-residency grounds — but the pure price argument no longer favours them in this tier.
Should I still keep GPT-5.6 Luna in my routing table?
Almost certainly not, unless you have a specific measured failure that only GPT-5.6 Luna passes. Its ceiling is one point higher and its coding index two points higher, but it costs 2.6× more per task on the independent measurement and carries a three-month-staler knowledge cutoff. The one thing it still has is a published system card, which GPT-6 Luna does not.
Bench test the swap before you route your volume
Take 200 items off your real queue, run them through both Luna generations, and score the outputs — not the tokens. That is the only way to see whether the regressions Artificial Analysis measured land on your traffic.
Open a new chat on CodingFleet →Sources & further reading
- Artificial Analysis — GPT-6 Sol and Luna push the cost efficiency frontier: cost per task, Coding Agent Index, AA-Omniscience, GDPval-AA and AA-Briefcase findings for both Luna generations.
- Artificial Analysis — GPT-5.6 Luna release page and the GPT-6 Luna release page: effort ladders, family ceilings and cost per task.
- OpenAI — Introducing GPT-6 Sol and Luna: AutomationBench, DeepSWE and OSWorld claims for Luna.
- VentureBeat — OpenAI releases GPT-6 Sol and Luna, slashing API costs 50% or more: pricing confirmation, the factuality evaluation, circumvention rates.
- Coursiv — GPT-6 Sol and Luna: pricing, benchmarks, availability: the full rate card including cache and long-context lines.
- ZDNET — OpenAI's GPT-6 Sol doubles its accuracy rate for half the cost.
- Microsoft Foundry — model retirement schedule: GPT-5.6 Luna retirement date and migration targets.
- Related on CodingFleet: GPT-6 Sol vs GPT-5.6 Sol, Claude Opus 5.5 vs GPT-6 Sol, GPT-6 Astra Review.
Benchmark scores are vendor- or evaluator-reported and labelled as such. Effort settings are stated wherever the source states them; where a vendor omits the effort level of a comparison, the article says so rather than assuming one. Cross-harness numbers are directional. Prices are USD per million tokens as of late September 2026.