OpenAI's cheapest tier has now had its price cut twice in eleven weeks. GPT-5.6 Luna launched at $1 / $6 per million tokens on 9 July, was cut to $0.20 / $1.20 on 30 July, and was replaced by GPT-6 Luna at $0.10 / $0.50 on 22 September. The same illustrative agent task that cost $3.75 in July costs $0.345 today — a 91% reduction in ten weeks.

That is not the whole story, and the rest of it is less flattering. Artificial Analysis ran both generations through the identical ten-evaluation harness, and GPT-6 Luna's ceiling went down.

TL;DR

GPT-6 Luna is a cost-and-safety release that gives up a little capability, and the independent numbers say so cleanly. Its top index score falls from 38 to 37 and its Coding Agent Index falls from 43 to 41, while cost per task drops 61% from $0.18 to $0.07. Hallucination falls from 93% to 77%, and its rate of working around an explicit "access denied" signal nearly halves.

  • The ceiling moved down, not up. 38 → 37 on Intelligence Index v4.3.2 — the first time this tier has lost a point generation-over-generation.
  • Coding regressed. Coding Agent Index 43 → 41, SWE-Atlas-QnA 49% → 44%, DeepSWE v1.1 66% → 64%.
  • Knowledge work regressed too. GDPval-AA v2.1 −75 Elo and AA-Briefcase v1.1 −45 Elo, traced by Artificial Analysis to weaker presentation and omitted rubric elements.
  • Safety improved sharply. Hallucination 93% → 77%, AA-Omniscience Index −10 → +1, restriction circumvention 76.5% → 42.4%.
  • It is now cheaper than the open-weight disruptors. At $0.10 / $0.50 it undercuts DeepSeek V4.1 Flash's off-peak rate of $0.15 / $0.60 and is 40× cheaper than GPT-5.6 Sol.
−61%
cost per task
$0.18 → $0.07
38 → 37
index ceiling
a point lower
93% → 77%
hallucination rate
AA-Omniscience
−91%
same task, ten weeks
$3.75 → $0.345

The quietest price collapse in the lineup

Luna's pricing is the part of this story that barely got covered, because Sol's headline cut was 50% and Luna's looked like more of the same. It wasn't. Luna has been repriced twice, and the cumulative change is an order of magnitude.

Line, per million tokensGPT-5.6 Luna
9 Jul 2026
GPT-5.6 Luna
after 30 Jul cut
GPT-6 Luna
22 Sep 2026
Total change
Input$1.00$0.20$0.10−90%
Output$6.00$1.20$0.50−92%
Cached input$0.10$0.02$0.01−90%
Cache write$1.25$0.25$0.125−90%
Over 272K input, in / out$2 / $9$0.40 / $1.50$0.20 / $0.75−90% / −92%
GPT-5.6 Luna · 9 July $6.00 GPT-5.6 Luna · after 30 July $1.20 input cut 80%, output cut 80% GPT-6 Luna · 22 September $0.50 input cut a further 50%, output a further 58% Ten-week total output −92% · input −90% — and unlike GPT-5.6 Sol, none of it is promotional
Output price per million tokens, three pricing points. Luna's cuts came from a permanent list-price change on 30 July, not a promotion — unlike Sol, whose $4/$20 rate carries a 21 November review date. Source: OpenAI API pricing and model pages.

The result is a model that has left the price band it used to occupy. Here is where GPT-6 Luna sits against the other cheap options teams actually compare it to.

ModelInput / MOutput / MNote
GPT-6 Luna$0.10$0.50Cached input $0.01 — a 90% read discount
GPT-5.6 Luna$0.20$1.20The model it replaces
DeepSeek V4.1 Flash$0.15 off-peak · $0.30 peak$0.60 off-peak · $1.20 peakUndercut on input and output at every hour
Claude Haiku 4.5$1.00$5.0010× more expensive on both lines
GPT-5.6 Sol$4.00$20.0040× more expensive on both lines
GPT-6 Sol$2.00$10.0020× more expensive; cache reads 20× the price
GPT-6 Astra$10.00$50.00100× more expensive on both lines

The competitive framing flipped. Through 2025 and most of 2026, the argument for hosted open-weight models was price: a frontier-adjacent model at a fraction of a flagship's rate. GPT-6 Luna at $0.10 / $0.50 now undercuts DeepSeek V4.1 Flash's off-peak rate, while carrying a 1.05M-token context window and six reasoning-effort levels. The price argument alone no longer favours open weights.

Spec sheet

SpecificationGPT-6 LunaGPT-5.6 LunaChange
Released22 September 20269 July 2026+75 days
API model IDgpt-6-lunagpt-5.6-luna the gpt-5.6 alias routes to Sol
Input / output$0.10 / $0.50$0.20 / $1.20−50% / −58%
Cached input$0.01$0.02−50%
Cache write$0.125$0.25−50%
Over 272K input, in / out$0.20 / $0.75$0.40 / $1.50−50%
Fast mode2× — $0.20 / $1.00not offered at this tiernew
Batch and Flex50% off — $0.05 / $0.2550% offunchanged structure
Context window1.05M tokens1.05M tokensunchanged
Max output128K tokens128K tokensunchanged
Knowledge cutoff18 May 202616 February 2026+3 months
Effort levelsnone, low, medium, high, xhigh, max — default mediumnone, low, medium, high, xhigh, maxunchanged
Where it runsOpenAI API, Codex, ChatGPT Work, and the desktop app for Free and Go usersAPI, Codex, ChatGPT Workwider
Five-hour message range, Plus350 – 3,000250 – 2,000+40% to +50%
System cardnone published at launchpublished with the GPT-5.6 familyregression

Note what did not change: the context window, the output ceiling, and the six-level effort ladder. GPT-6 Luna is the same shape as GPT-5.6 Luna with a three-month-fresher knowledge cutoff, a tenth of the price, and — as the next section shows — a slightly lower top end.

The headline nobody wrote: the ceiling went down

Every previous release in this line has moved up or held. GPT-6 Luna moved down by a point.

Metric at max effortGPT-6 LunaGPT-5.6 LunaDelta
Artificial Analysis Intelligence Indexv4.3.2, family ceiling3738−1
Cost per Intelligence Index task$0.07$0.18−61%
Output tokens per taskthe model works harder for less51k41k+24%
AA Coding Agent IndexCodex harness4143−2
SWE-Atlas-QnAcodebase question answering44%49%−5
DeepSWE v1.1long-horizon software engineering64%66%−2
AutomationBench-AAbusiness workflows · partial credit53%50%+3
Terminal-Bench 4.0agentic terminal work13%12%+1
GDPval-AA v2.1knowledge work · Elo−75 Elo
AA-Briefcase v1.1work documents · Elo−45 Elo

Cost per task fell 61% while token use rose 24%. That gap is the whole economics of the release: 51k output tokens a task against 41k, and still less than half the bill, because the rate card dropped faster than the workload grew. It also means the price cut is doing all the work — there is no efficiency gain here, and slightly an efficiency loss.

The two knowledge-work regressions deserve the same caution Artificial Analysis gives them: the drop "tends to be driven by reduced presentation quality and deliverables that omit rubric elements," and their team "has manually inspected hundreds of model outputs" to establish that. On a classification or extraction job it is irrelevant. On a client deliverable with a required structure, it is the entire output.

Where GPT-6 Luna is a clear improvement: safety

The safety story is the trade the price cut bought, and on three separate measures it is a large step rather than a marginal one. All three come from Artificial Analysis's independent runs.

GPT-6 Luna · max GPT-5.6 Luna · max Hallucination rate — lower is better 77% 93% Worked around an "access denied" signal 42.4% 76.5% Accuracy on questions attempted 44% 43%
Two large wins and one flat line. Hallucinations fall 16 points and adversarial circumvention falls 34 points, while accuracy holds at 43–44%. On the composite AA-Omniscience Index — which rewards correct answers and penalises hallucinations, with no penalty for abstaining — GPT-6 Luna moves from −10 to +1: from net-negative to net-positive knowledge reliability. Source: Artificial Analysis.

The circumvention number is the striking one. In adversarial runs containing an explicit "access denied" signal, GPT-5.6 Luna attempted to work around the restriction 76.5% of the time. GPT-6 Luna does so 42.4% of the time. OpenAI stresses that these are deliberately difficult, low-stakes tests run without the system-level safeguards its products apply, so they should not be read as normal-use failure rates — but a 34-point move in the same test is a real behavioural change, and it is the largest single improvement in the release.

Why this matters more than the index point. A cheap model is only useful in unattended volume work if you can predict how it fails. GPT-6 Luna still hallucinates on 77% of the questions it gets wrong, which is high in absolute terms — but it is now a model whose errors are less likely to be confident inventions and less likely to be boundary-pushing workarounds. For a tier designed to be called thousands of times a day, that is worth more than one index point.

The crossover that matters: GPT-6 Luna vs GPT-5.6 Sol

The comparison that decides most routing questions is not Luna against Luna. It is OpenAI's new cheap tier against its old mid tier — because both are now options on the same pricing page, and the cheap one is 40× cheaper per token.

MeasureGPT-6 LunaGPT-5.6 SolVerdict
Input / output per million$0.10 / $0.50$4 / $2040× cheaper
Cached input$0.01$0.4040× cheaper
Intelligence Index ceiling374710 points behind
Knowledge cutoff18 May 202616 February 20263 months fresher
Context window1.05M1.05Mtie
OSWorld 2.0 OpenAI's own runLuna (max) exceeds Sol (medium) at ~1/10 the costLuna
Factuality OpenAI's internal evalMatches at higher effort for ~1/100 the task costLuna
DeepSWE v1.1 OpenAI's run66.6%72.7% at GPT-5.6 Sol (max)6.1 points behind

OpenAI's own framing for Luna's DeepSWE result is that 66.6% is "comparable to Claude Opus 5 and Fable 5 at medium effort," at 93% less cost per task than Opus 5 and 96% less than Fable 5. Read next to the row above it, the picture is consistent: Luna is not a replacement for a mid-tier model on hard problems, but on the tasks where Luna is competent it now beats last generation's mid tier on cost by an order of magnitude, and on computer use it beats it outright.

That is the real routing question for anyone running a tiered pipeline. If your work is inside GPT-5.6 Luna's competence band and you validate outputs downstream, GPT-6 Luna is a straight swap that cuts your bill by another 61%. If your work needed GPT-5.6 Sol, GPT-6 Luna will not cover it — and the regressions in this release mean you should re-test rather than assume.

What the price cut is worth on a real invoice

Here is the illustrative cache-heavy agent task — 8M cache reads, 400K uncached input, 600K cache writes, 300K output — at all three of Luna's historic rate cards. Token counts are held identical so only the published rates move.

LineGPT-6 LunaGPT-5.6 Luna (30 Jul)GPT-5.6 Luna (9 Jul)
Cache reads · 8,000,000$0.08$0.16$0.80
Uncached input · 400,000$0.04$0.08$0.40
Cache writes · 600,000$0.075$0.150$0.750
Output · 300,000$0.150$0.360$1.800
Total$0.345$0.750$3.750
Versus GPT-6 Luna2.2× more10.9× more

Two things to take from that table. First, the ten-week reduction is 91% on the same workload — a bigger drop than any single release in this family has delivered. Second, the measured independent figure is a 61% cut rather than 54%, because GPT-6 Luna's real workloads use more tokens than its predecessor's, and the token growth partially offsets the rate cut in a way this fixed-token illustration does not capture.

Verdict

The short version

If your work already fits inside GPT-5.6 Luna's competence band, switch — the bill falls 61% and the failure modes get safer. Hallucinations are down 16 points, adversarial circumvention is down 34 points, the knowledge cutoff is three months fresher, and the rate card has no promotional expiry date on any line.

If you were using Luna as a cheap coding agent, be more careful. The Coding Agent Index fell two points, SWE-Atlas-QnA fell five, and both knowledge-work Elo measures fell. That is a real capability regression in exchange for a real price cut, and whether it is worth it depends on where your tasks sit relative to the ceiling. Test the hard 10% of your queue before you migrate the whole thing.

Your workloadRoute toWhy
High-volume classification, extraction, tagging, routingGPT-6 Luna61% cheaper per task with better hallucination behaviour — the tier's core use case
First-pass drafting before an escalation gateGPT-6 LunaSafer failure mode at a tenth of the tier above it
Business workflow automationGPT-6 LunaAutomationBench-AA improved 50% → 53%; OpenAI reports +5.4 points at 58% lower cost
Anything involving guardrails and adversarial inputsGPT-6 LunaCircumvention halved, from 76.5% to 42.4%
Volume work inside an open-weight budgetGPT-6 LunaUndercuts DeepSeek V4.1 Flash's off-peak rate at a 1.05M context window
Cheap coding agentsTest firstCoding Agent Index −2, SWE-Atlas-QnA −5, DeepSWE −2
Long documents with required structureTest firstGDPval-AA −75 Elo, AA-Briefcase −45 Elo on presentation and omitted rubric items
Anything where the ceiling mattersNeither LunaGPT-5.6 Sol is 10 index points ahead; GPT-6 Sol is 11

FAQ

Is GPT-6 Luna worse than GPT-5.6 Luna?

On capability at the top end, yes, slightly: the Intelligence Index ceiling falls from 38 to 37 and the Coding Agent Index from 43 to 41. On cost and safety it is much better: 61% cheaper per task, hallucination 93% → 77%, and restriction circumvention 76.5% → 42.4%. It is a different trade, not a straight downgrade — and for high-volume work inside its band the trade is favourable.

Why did OpenAI cut Luna's price so far?

OpenAI says "improvements to inference and caching" let it reduce prices while increasing capability, and it confirmed to VentureBeat that the GPT-6 rates are permanent rather than promotional. The 30 July cut that took Luna from $1/$6 to $0.20/$1.20 was also a permanent list-price change — unlike Sol's August discount, which has a 21 November review date. Note that capability and price moved in opposite directions here: Sol improved while holding price, Luna cut price while slipping a point.

Does the lower hallucination rate mean it is more accurate?

No — accuracy is flat at 43–44% on questions attempted. What improved is that Luna produces a wrong answer far less often, because it hallucinates less. On the composite AA-Omniscience Index, which rewards correct answers and penalises hallucinations without penalising abstention, GPT-6 Luna moves from −10 to +1. That is a move from net-negative to net-positive knowledge reliability.

Is GPT-6 Luna cheaper than open-weight models now?

On the hosted rates teams compare, yes. At $0.10 input / $0.50 output it undercuts DeepSeek V4.1 Flash's off-peak rate of $0.15 / $0.60, and it is roughly 10× cheaper than Claude Haiku 4.5 at $1 / $5. Open weights still win on self-hosting, fine-tuning and data-residency grounds — but the pure price argument no longer favours them in this tier.

Should I still keep GPT-5.6 Luna in my routing table?

Almost certainly not, unless you have a specific measured failure that only GPT-5.6 Luna passes. Its ceiling is one point higher and its coding index two points higher, but it costs 2.6× more per task on the independent measurement and carries a three-month-staler knowledge cutoff. The one thing it still has is a published system card, which GPT-6 Luna does not.

Bench test the swap before you route your volume

Take 200 items off your real queue, run them through both Luna generations, and score the outputs — not the tokens. That is the only way to see whether the regressions Artificial Analysis measured land on your traffic.

Open a new chat on CodingFleet →

Sources & further reading

Benchmark scores are vendor- or evaluator-reported and labelled as such. Effort settings are stated wherever the source states them; where a vendor omits the effort level of a comparison, the article says so rather than assuming one. Cross-harness numbers are directional. Prices are USD per million tokens as of late September 2026.