On 22 September, OpenAI launched GPT-6 Luna at $0.10 / $0.50 per million tokens — and quietly made the cheapest model in its own catalogue obsolete. GPT-5.4 mini, the March release that was supposed to be the budget workhorse, still lists at $0.75 / $4.50. On a representative 10M-input + 1M-output workload, that is $1.50 versus $12.00 — 8× cheaper for the newer model, with 2.6× the context window and a knowledge cutoff eight and a half months fresher.
So is GPT-5.4 mini simply finished? Not quite, and the reason is unusual: OpenAI published almost no shared evaluations between the two models. GPT-5.4 mini's scorecard is built on SWE-Bench Pro, GPQA and Toolathlon. GPT-6 Luna's is built on DeepSWE, OSWorld 2.0 and AutomationBench. The overlap is essentially one near-miss on the OSWorld family, run under two different protocols. This article separates what is actually measured from what is marketing adjacency — and shows where each model still earns its place.
TL;DR
On every directly comparable line, GPT-6 Luna wins: price (8× cheaper on a real mix), context (1.05M vs 400K), knowledge cutoff (May 2026 vs August 2025), and the effort ladder (six levels including max, versus mini's five). On the one harness that has measured both — Artificial Analysis's Intelligence Index — Luna (max) scores 37 to mini's 24 (xhigh) at 6.4× lower cost per task. Mini's counter-case is speed (~224 output tokens/s, the fastest thing in its price tier) and delegation: OpenAI still positions it for subagents and fast interactive loops.
- The bill: 10M input + 1M output costs $1.50 on GPT-6 Luna vs $12.00 on GPT-5.4 mini at list prices — and $0.27 vs $2.25 on a cache-heavy mix.
- The shared independent numbers favour Luna. On AA's harness: Intelligence Index 37 vs 24, cost per index task $0.07 vs $0.45, accuracy 44% vs 37.5%, hallucination-on-error 77% vs 90%. But mini is ~40% faster token-for-token, and AA has deprecated mini's coverage, so its figures are historical.
- OpenAI's own scorecards still share nothing. Vendor-side, mini publishes SWE-Bench Pro/GPQA/MCP Atlas rows and Luna publishes DeepSWE/OSWorld/AutomationBench rows — the two tables never overlap. The OSWorld near-miss (72.1% Verified vs 52.7% 2.0 offline) uses two different protocols.
- Mini is not retired. It remains on the API at list price, keeps free-tier ChatGPT access, and remains the designated subagent model in OpenAI's own guidance — though AA now redirects mini readers to GPT-5.6 Terra.
$1.50 vs $12.00
2.6× the capacity
May 2026 vs Aug 2025
same harness, both models
The rate card: every line favours Luna
GPT-5.4 mini launched in March at what was then a sharp price. GPT-6 Luna's September launch lands below it on every line of the rate card — input, output, cached input, and the over-272K long-context tier — while carrying a larger window. For context, GPT-5.6 Luna, the intermediate generation, sits between the two.
| Per 1M tokens | GPT-6 Luna | GPT-5.4 mini | GPT-5.6 Luna (ref) |
|---|---|---|---|
| Input | $0.10 | $0.75 | $0.20 |
| Output | $0.50 | $4.50 | $1.20 |
| Cached input | $0.01 | $0.075 | $0.02 |
| Over 272K input, in / out | $0.20 / $0.75 | no separate tier | $0.40 / $1.50 |
| Batch / Flex | 50% off — $0.05 / $0.25 | available | 50% off |
| Fast mode | 2× — $0.20 / $1.00 | — | — |
The cache-heavy case is just as decisive. Take 8M cached reads, 400K fresh input and 300K output — a typical RAG or long-conversation agent profile:
| Line | GPT-6 Luna | GPT-5.4 mini |
|---|---|---|
| Cached reads · 8,000,000 | $0.08 | $0.60 |
| Fresh input · 400,000 | $0.04 | $0.30 |
| Output · 300,000 | $0.15 | $1.35 |
| Total | $0.27 | $2.25 |
| Versus GPT-6 Luna | — | 8.3× more |
One caveat on third-party pricing: some resellers now offer GPT-5.4 mini at $0.375 / $2.25 — half the list rate. Even at those discounted rates, GPT-6 Luna's list price still wins on every line, and Luna's Batch/Flex tier ($0.05 / $0.25) undercuts mini's discounted output price. The comparison above uses each model's official OpenAI list price.
Spec sheet
| Specification | GPT-6 Luna | GPT-5.4 mini |
|---|---|---|
| Released | 22 September 2026 | 17 March 2026 |
| API model ID | gpt-6-luna | gpt-5.4-mini / gpt-5.4-mini-2026-03-17 |
| Positioning | Cheapest GPT-6 tier — high-volume focused work | GPT-5.4 small model for coding, computer use and subagents |
| Context window | 1.05M tokens | 400K tokens |
| Max output | 128K tokens | 128K tokens |
| Knowledge cutoff | 18 May 2026 | 31 August 2025 |
| Input / output | Text and images / text | Text and images / text |
| Reasoning effort | none, low, medium, high, xhigh, max | none, low, medium, high, xhigh |
| Tools | Function calling, web search, file search, computer use | Web, file, code, shell, computer use, MCP, tool search, skills |
| Output speed | 140–163 tok/s AA measured | ≈224 tok/s AA measured |
| Availability | API, Codex, ChatGPT Work; desktop app for Free and Go users | API, Codex; limited free-tier ChatGPT access |
| License | Proprietary API | Proprietary API |
Two structural differences matter beyond the rate card. First, context: Luna's 1.05M window holds an entire mid-sized repository or a long research trail without retrieval scaffolding; mini's 400K requires compaction far earlier. Second, the effort ladder: Luna's max level is where its best published numbers live (DeepSWE 66.6%, OSWorld 2.0 52.7%), and it simply has no equivalent on mini. Mini counters with a broader tool surface — it is the only one of the two that ships OpenAI's skills and tool-search layer.
The comparison problem: they publish on almost nothing shared
This is the section most comparisons skip, and it is the most important one. OpenAI's evaluation reporting changed between the GPT-5.4 generation (March) and the GPT-6 generation (September). The result: zero directly comparable rows on the vendors' own scorecards. (Artificial Analysis's shared harness does measure both models — that comes in the next section. This section is about what OpenAI itself publishes.)
| Published rows — GPT-6 Luna only | Score | Conditions |
|---|---|---|
| DeepSWE 1.1 long-horizon software engineering | 66.6% | max effort, OpenAI's own run |
| OSWorld 2.0 offline computer use | 52.7% | max effort; ladder: low 8.3% → medium 31.5% → high 41.4% → xhigh 46.7% |
| AutomationBench-AA business workflows | 53% | Artificial Analysis harness |
| SWE-Atlas-QnA codebase Q&A | 44% | AA harness, max |
| Terminal-Bench 4.0 public board | 16.4% ±2.7 | Codex harness, $0.1k run cost |
| Published rows — GPT-5.4 mini only | Score | Conditions |
|---|---|---|
| SWE-Bench Pro public split | 54.4% | xhigh effort |
| GPQA Diamond | 88.0% | OpenAI scorecard |
| Toolathlon | 42.9% | OpenAI scorecard |
| MMMU Pro, no tools | 76.6% | OpenAI scorecard |
| OSWorld-Verified computer use | 72.1% | different protocol from OSWorld 2.0 |
| MCP Atlas tool orchestration | 57.7% | OpenAI scorecard |
| HLE with tools | 41.5% | OpenAI scorecard |
| Terminal-Bench 2.0 | 60.0% | older task set — not comparable to TB 4.0 |
Read those two tables as an honest gap, not a scoreboard. GPT-6 Luna has no published SWE-Bench Pro, GPQA or MMMU number; GPT-5.4 mini has no DeepSWE or AutomationBench counterpart. The only near-miss is computer use:
The shared harness: Artificial Analysis has measured both
Artificial Analysis runs every current OpenAI model through its identical evaluation battery, which makes it the only harness covering both of these models. For GPT-6 Luna, those numbers — published in our GPT-6 Luna vs GPT-5.6 Luna comparison — set the ceiling conversation: 37 on the Intelligence Index at max effort, one point below GPT-5.6 Luna, with a Coding Agent Index of 41 (−2) and a measured $0.07 per index task — a 61% cost reduction that comes from the rate card, not from working less (output tokens per task actually rose 24%, from 41k to 51k). On safety behaviour, hallucination at max effort fell from 93% to 77%, and adversarial circumvention of an "access denied" signal nearly halved (76.5% → 42.4%).
Artificial Analysis has also measured GPT-5.4 mini — at its top xhigh effort, since the model has no max level. Here are the rows where the two can be placed side by side:
| AA measurement (same harness) | GPT-6 Luna (max) | GPT-5.4 mini (xhigh) | Reading |
|---|---|---|---|
| Intelligence Index | 37 | 24 | Luna +13 — and mini has no higher effort level to close the gap |
| Cost per Intelligence Index task | $0.07 | $0.45 | Luna 6.4× cheaper per unit of measured work |
| AA-Omniscience accuracy | 44% | 37.5% | Luna answers more questions correctly |
| AA-Omniscience non-hallucination rate | 23% | 9.8% | When wrong, Luna hallucinates 77% of the time; mini 90.2% |
| Output speed | 140–163 tok/s | ≈224 tok/s | Mini is genuinely faster — its strongest shared-harness result |
| Coding index | 41 agent-based (Codex) | 56.1 model-level | Different constructs — listed, not compared |
| Terminal-Bench family | 16.4% on TB 4.0 | 52.3% on TB Hard | Different task sets and years — listed, not compared |
Two caveats, one in each direction. First, Artificial Analysis has marked GPT-5.4 mini as deprecated — it now only refreshes the default 10k-token workload and points mini readers to GPT-5.6 Terra — so mini's index numbers are historical rather than wrong, and its broader AA coverage is frozen (GPQA Diamond 87.5%, AA-LCR 77.0%, τ²-Bench Telecom 83.3%, SciCode 52.1%). Second, the coding indexes are not the same construct: Luna's 41 measures a model-plus-harness agent, mini's 56.1 measures the model alone. We report both without equating them. What survives every caveat is the shape of the result: on the harness that measures both, Luna is 13 index points ahead at 6.4× lower cost, and mini's remaining edge is speed.
Terminal-Bench 4.0: the one place both generations face the same board
The public tbench.ai Terminal-Bench 4.0 board gives the cleanest like-for-like data point in this article — not versus mini, which never ran it, but versus the tier Luna replaced:
| Model | Resolution rate | Run cost | Source |
|---|---|---|---|
| GPT-5.6 Luna | 17.3% | — | public board row |
| GPT-6 Luna | 16.4% ±2.7 | $0.1k | public board row · max effort · Codex |
| GPT-5.4 mini | not on the board | — | its 60.0% row is Terminal-Bench 2.0, a different, easier task set |
Note the direction: on the hardest public terminal board, GPT-6 Luna scores a point below the GPT-5.6 Luna it replaces — at a fraction of the run cost. That pattern (capability flat-to-slightly-down, price down hard) is the signature of this release, and it is exactly the trade a budget tier should make. But it is also a warning against assuming "new generation" means "better" here.
Where GPT-5.4 mini still earns its place
| Mini's remaining case | Why it still holds |
|---|---|
| Latency and delegation | OpenAI's own positioning: mini is the fast worker for subagents, screenshot interpretation and high-volume loops, running more than twice as fast as GPT-5 mini at 30% of GPT-5.4's Codex quota. Luna is a reasoning-tier model; its max numbers come with 51k output tokens per task. |
| Published breadth | Nine-plus scorecard rows (SWE-Bench Pro 54.4%, MCP Atlas 57.7%, HLE with tools 41.5%) plus frozen AA-harness coverage (GPQA 87.5%, τ²-Bench 83.3%, AA-LCR 77.0%, SciCode 52.1%) in categories where Luna publishes little — the caveat being that AA has deprecated mini, so those rows are historical. |
| Tool surface | Skills, tool search and the full GPT-5.4 agent stack. Luna's built-in set (web, file search, computer use) is narrower on paper. |
| Free-tier presence | Mini retains limited free-tier ChatGPT access; it remains the default "cheap OpenAI model" in many SDK templates and tutorials. |
None of these survive a pure price-per-token comparison — but not every workload is priced per token. A subagent that answers in 400 tokens at low effort cares about time-to-first-token and message-per-hour limits, and that is the workload mini was shaped for.
Routing guide
| Your workload | Route to | Why |
|---|---|---|
| High-volume classification, extraction, tagging | GPT-6 Luna | 8× cheaper on token rates, larger context, fresher cutoff |
| Repository-scale context, long research bundles | GPT-6 Luna | 1.05M vs 400K — no compaction scaffolding needed |
| Budget agentic coding with reasoning headroom | GPT-6 Luna | DeepSWE 66.6% at max; six-level effort ladder; $0.07/task measured cost |
| Fast subagents doing bounded supporting work | GPT-5.4 mini | OpenAI's designated delegation model; tuned for short feedback loops |
| Latency-sensitive interactive product traffic | GPT-5.4 mini | ≈224 tok/s measured vs Luna's 140–163 — roughly 40% faster token-for-token; still test TTFT on your own traffic |
| Screenshot-driven computer use | Test both | Mini's 72.1% (OSWorld-Verified) and Luna's 52.7% (OSWorld 2.0 offline) are different protocols — the honest answer is a shadow test |
| Structured client deliverables | Test first | Luna's knowledge-work Elo regressed vs GPT-5.6 Luna on presentation quality; mini's equivalent is unmeasured |
| Anything needing a documented scorecard row | Check both | The two models publish disjoint evaluation sets — see the tables above |
Verdict
The short version
For new builds, start with GPT-6 Luna. It is 8× cheaper than GPT-5.4 mini at list prices on both fresh and cached lines, leads the shared Artificial Analysis harness 37 to 24 on the Intelligence Index at 6.4× lower cost per task, carries 2.6× the context, a May 2026 knowledge cutoff, and better hallucination behaviour on the same AA-Omniscience runs. Every directly comparable measurement favours it. There is no pricing scenario in which mini wins.
Keep GPT-5.4 mini for the latency tier. It remains the measurably faster model (~224 tok/s vs 140–163, the one shared row it wins) and the delegation-oriented one, with a broader frozen scorecard and the best-documented computer-use result of the pair (72.1% OSWorld-Verified). But treat it as a specialist now, not a default: if your mini workload is volume-driven rather than latency-driven, run a 200-item shadow test on Luna before paying mini's 8× premium — and note that even Artificial Analysis now redirects mini readers to a newer model.
FAQ
Is GPT-6 Luna simply better than GPT-5.4 mini?
On everything directly measurable, yes: price (8× at list rates), context (1.05M vs 400K), knowledge freshness (May vs August 2025), and — on Artificial Analysis's shared harness — capability, where Luna scores 37 vs mini's 24 on the Intelligence Index, answers more questions correctly (44% vs 37.5% accuracy), and hallucinates far less when wrong (77% vs 90%). The caveats: OpenAI's own scorecards still share no rows, and mini's AA numbers are historical (deprecated coverage) rather than freshly measured. Mini's one clear win on the shared harness is speed: ~224 tok/s vs Luna's 140–163.
Should I migrate my GPT-5.4 mini workload to GPT-6 Luna?
If the workload is high-volume and latency-tolerant — classification, extraction, summarisation, budget coding — yes, and expect roughly an 87% bill reduction plus a 13-point index lead. If it is latency-sensitive (interactive chat, tight subagent loops), benchmark time-to-first-token and throughput on both first, since mini's ~224 tok/s is its strongest measured advantage. And re-test the hard tail of your queue: the shared index says Luna is well ahead, but your traffic is the final word.
Why does GPT-5.4 mini still cost $0.75 / $4.50?
OpenAI has not repriced it since March. Some third-party resellers offer it at $0.375 / $2.25, but even that does not close the gap with Luna's $0.10 / $0.50 list rate. Given that OpenAI cut GPT-5.6 Luna's prices twice in eleven weeks before replacing it, a mini repricing or retirement notice is plausible — but nothing has been announced, and it still lists as active.
Which model is better for computer use?
Genuinely unclear, and the published numbers must not be mixed: mini's 72.1% is on OSWorld-Verified (older protocol), Luna's 52.7% is on OSWorld 2.0 offline (newer, harder protocol). What is clear is OpenAI's own cost framing: Luna (max) beats GPT-5.6 Sol (medium) on OSWorld 2.0 at one-tenth the cost. For production computer-use work, test screenshot resolution, click policy and retry budget on both.
How does GPT-6 Luna compare to the rest of the cheap tier?
At $0.10 / $0.50 it undercuts DeepSeek V4.1 Flash's off-peak rate ($0.15 / $0.60) and is roughly 10× cheaper than Claude Haiku 4.5 ($1 / $5) — with a 1.05M context window that neither matches. See our GPT-6 Luna vs GPT-5.6 Luna comparison for the full competitive picture and the safety-behaviour data.
Bench test the swap before you route your volume
Take 200 items off your real queue, run them through GPT-6 Luna and GPT-5.4 mini, and score the outputs — not the tokens. The shared index says Luna is 13 points ahead; whether that advantage lands on your traffic is a question only your own eval can answer.
Open a new chat on CodingFleet →Sources & further reading
- OpenAI — Introducing GPT-6 Sol and Luna: Luna's rate card, DeepSWE 66.6%, OSWorld 2.0 cost framing, AutomationBench.
- OpenAI — Introducing GPT-5.4 mini and nano: mini's scorecard rows, subagent positioning, availability.
- Artificial Analysis — GPT-6 Sol and Luna push the cost-efficiency frontier: Intelligence Index 37, Coding Agent Index 41, $0.07/task, hallucination and circumvention rates.
- Artificial Analysis — GPT-5.4 mini (xhigh) model page and release page: Intelligence Index 24, $0.45/task, ≈224 tok/s, AA-harness GPQA/τ²-Bench/SciCode rows, and the deprecation notice.
- tbench.ai — Terminal-Bench 4.0 public leaderboard: GPT-6 Luna 16.4% ±2.7 ($0.1k run) and GPT-5.6 Luna 17.3% public rows.
- GPT-6 Luna API documentation and GPT-5.4 mini API documentation: context, modalities, tools, effort levels, token rates.
- BenchLM — GPT-6 Luna profile and GPT-5.4 mini profile: spec and pricing cross-checks.
- Related on CodingFleet: GPT-6 Luna vs GPT-5.6 Luna, GPT-5.6 Luna vs GPT-5.4 Mini, Terminal-Bench 4.0 Leaderboard.
Benchmark scores are labelled by source: "OpenAI scorecard" rows are vendor-reported at the stated effort level; Artificial Analysis rows are independent runs in its own harness; tbench.ai rows come from the public board's Codex harness. Prices are USD per million tokens at OpenAI list rates as of October 2026. Where two models publish on different evaluations, this article reports the gap instead of manufacturing a winner.