DeepSeek V4.1 Flash (DeepSeek, September 10, 2026) and Claude Opus 5 (Anthropic, July 24, 2026) sit at the two ends of the frontier's price curve. On DeepSeek's own published table, V4.1 Flash — a 552B-parameter model that activates just 8B parameters on input and 16B on output — matches or beats a flagship that costs 16.7× to 41.7× more per token on five of the fifteen benchmarks where both have numbers.
But "matches" hides a lot. DeepSeek's five wins average 1.94 points. Opus 5's ten wins average 10.19 points. One model wins short sprints by a nose; the other wins long problems by a mile. That distinction, not the headline scores, is the whole comparison.
Everything below comes from DeepSeek's release notes and pricing page, Anthropic's Opus 5 launch materials, Artificial Analysis, and hands-on reports from people who ran V4.1 Flash during its two-day beta. Harnesses differ between labs, so vendor rows are labelled as vendor rows.
At a Glance
| Category | DeepSeek V4.1 Flash | Claude Opus 5 |
|---|---|---|
| Released | September 10, 2026 (beta from Sept 8) | July 24, 2026 |
| Developer | DeepSeek | Anthropic |
| Architecture | Causal Encoder–Decoder (CED) MoE, 552B total — 8B active on input, 16B on output | Opus-class flagship (undisclosed size) |
| Context window | 1M tokens | 1M tokens |
| Max output | 384K tokens | 128K tokens |
| Multimodal input | Yes — native vision trained jointly from pre-training | Yes — image input |
| Reasoning control | Thinking (default) + non-thinking modes | Adaptive thinking, low → max effort |
| API surface | OpenAI-compatible, Anthropic format, Responses API, tool calls, JSON mode, FIM (non-thinking) | Claude API, mid-conversation tool changes (beta), automatic fallbacks (beta) |
| Input price | $0.15 off-peak / $0.30 peak per 1M | $5.00 per 1M |
| Cached input | $0.003 off-peak / $0.006 peak per 1M | $0.50 per 1M |
| Output price | $0.60 off-peak / $1.20 peak per 1M | $25.00 per 1M |
| Fast mode | Not offered (the whole model is the fast tier) | ~2.5× speed at 2× price |
| Concurrency ceiling | 2,500 concurrent requests | Not published as a single figure |
| Model ID | deepseek-flash (legacy deepseek-v4-flash routed to it) | claude-opus-5 |
| Weights | Published on Hugging Face with a technical report | Closed |
Sources: DeepSeek API docs (changelog, Models & Pricing, launch note), Anthropic's Claude Opus 5 launch post, OpenRouter model page. DeepSeek's own family (V4-Flash and V4-Flash-Vision-Exp) has been retired and now routes to V4.1 Flash.
The Vendor Benchmark Table, Head to Head
DeepSeek shipped a 19-benchmark table for V4.1 Flash that includes Claude Opus 5 as an external comparison point. Here it is reduced to the two model columns, with each row marked by whoever leads.
| Benchmark | DeepSeek V4.1 Flash | Claude Opus 5 | Leader |
|---|---|---|---|
| GPQA Diamond | 90.9 | 93.4 | Opus 5 |
| HLE | 36.8 (39.1*) | 56.3 | Opus 5 |
| Codeforces (rating) | 3471 | — | V4.1 Flash |
| MathArena Apex | 65.6 | — | V4.1 Flash |
| Terminal-Bench 2.1 | 90.6 | 89.1 | V4.1 Flash (+1.5) |
| Terminal-Bench 3.0 | 30.0 | 43.3 | Opus 5 |
| Terminal-Bench 4.0 | 31.2 | 51.8 | Opus 5 |
| DeepSWE v1.1 | 74.2 | 74.0 | V4.1 Flash (+0.2) |
| ProgramBench | 20.3 | 37.0 | Opus 5 |
| NL2Repo-Bench | 65.4 | 75.3 | Opus 5 |
| CyberGym | 88.1 | — | V4.1 Flash |
| SEC-Bench Pro | 62.8 | — | V4.1 Flash |
| ExploitGym | 15.3 | 22.1 | Opus 5 |
| HLE (w/tools) | 63.9 | 63.6 | V4.1 Flash (+0.3) |
| Automation-Bench | 54.8 | 50.3 | V4.1 Flash (+4.5) |
| Agents' Last Exam | 31.8 | 28.6 | V4.1 Flash (+3.2) |
| Chartography (w/tools) | 78.9 | 84.0 | Opus 5 |
| BabyVision (w/tools) | 89.6 | 94.1 | Opus 5 |
| ZeroBench-main (w/tools) | 49.0 | 52.0 | Opus 5 |
Source: DeepSeek API changelog (2026-09-10) and the DeepSeek-V4.1 Flash launch note. * denotes the text-only subset of HLE. Rows marked "—" are benchmarks Anthropic's model was not reported on in DeepSeek's table, not losses. Vendor-reported; harnesses and effort settings differ across labs.
Coding & Terminal Agents
The pattern is clean. On the two newest terminal suites — Terminal-Bench 3.0 and 4.0, which target frontier, long-horizon CLI work — Opus 5 leads by 13.3 and 20.6 points. On the older, more widely reproduced Terminal-Bench 2.1, V4.1 Flash is ahead by 1.5. Read those together and the story is about task difficulty rather than model generation: the further the suite pushes past conventional agent loops, the wider Opus 5's margin gets.
DeepSWE v1.1 — the long-horizon software engineering suite we track on our own DeepSWE v1.1 leaderboard — is the closest thing to a genuine tie in the table: 74.2 vs 74.0. Worth noting that our leaderboard lists Opus 5 at 68.8% under a different harness, which is a good reminder that a two-point vendor-run gap is well inside methodology noise.
Reasoning, Knowledge & Vision
This is Opus 5's home turf, and the margins are not close. HLE: 56.3 vs 36.8 — a 19.5-point gap on the hardest published knowledge benchmark. ProgramBench: 37.0 vs 20.3 and NL2Repo: 75.3 vs 65.4. And on the three tool-augmented vision suites (Chartography, BabyVision, ZeroBench), Opus 5 sweeps, which is notable because V4.1 Flash is the first DeepSeek Flash built with native multimodal training rather than a bolted-on vision encoder.
Agentic, Business & Security Tasks
Here DeepSeek takes four of six, including the agentic-security rows where Opus 5 has no published number in DeepSeek's table at all. That gap is deliberate on Anthropic's side: Opus 5 was intentionally not trained on cyber tasks, and Anthropic states plainly that it trails Mythos 5 on exploit development. Its safeguard classifiers block binary-based vulnerability scanning, penetration testing, and exploit generation. If your workload is offensive security, that is a policy difference, not a benchmark difference.
Capability Radar
Normalized 0–100 for visual comparison only. Reasoning & knowledge = mean of GPQA Diamond and HLE. Frontier terminal = mean of Terminal-Bench 3.0 and 4.0. Coding agents = mean of DeepSWE v1.1 and NL2Repo. Tool-use & automation = mean of HLE (w/tools), Automation-Bench and Agents' Last Exam. Vision & charts = mean of Chartography, BabyVision and ZeroBench. Cost efficiency = cheapest peak output price divided by each model's output price × 100 ($1.20 ÷ $25 × 100 = 4.8 for Opus 5).
What DeepSeek Actually Changed
V4.1 Flash is not a post-training refresh of V4 Flash — and the difference matters if you had prompts tuned to the old model. DeepSeek describes a new Causal Encoder–Decoder architecture with an asymmetric split: 8B parameters active for input, 16B for output, out of a 552B-parameter MoE. Both figures come from DeepSeek's launch note. The company frames the design as built for "a higher capability ceiling, faster inference, higher throughput, and scaling to larger models."
The second change is cheaper agent loops. DeepSeek says V4.1 Flash's KV cache needs 1/4 the HBM and 1/8 the SSD storage of the previous generation. On agent loops that re-read a large context every turn, cache-hit tokens are most of the input bill — so compressing the cache compresses the recipe that made DeepSeek cheap in the first place.
Third, DeepSeek is retiring its own flagship in favor of it. From 04:00 UTC on September 14, 2026 — through the eventual launch of V4.1 Pro — every request to deepseek-v4-pro is served by V4.1 Flash and billed at V4.1 Flash prices. On DeepSeek's own table, V4.1 Flash beats V4-Pro 0813 on Terminal-Bench 2.1 (90.6 vs 87.9), DeepSWE (74.2 vs 62.7), CyberGym (88.1 vs 83.3), SEC-Bench Pro (62.8 vs 56.4), ExploitGym (15.3 vs 5.4), HLE with tools (63.9 vs 60.0), Automation-Bench (54.8 vs 43.2) and Agents' Last Exam (31.8 vs 25.7) — while still losing on GPQA Diamond (90.9 vs 92.4) and HLE text-only (36.8 vs 42.7). A vendor routing premium traffic to its own cheaper model is the strongest endorsement in the release.
What People Built and Said
V4.1 Flash had an unusual launch: a two-day API beta under the model ID deepseek-v4.1-flash-expires-on-0910, capped at 20 concurrent requests, before the formal September 10 release. That produced a burst of first-day testing — and the reports are consistent about what the model is good at, and consistent about where it wobbles.
- Speed is the headline everyone agrees on. World of AI measured peaks around 427 tokens/second, with sustained output near 400 tok/s, and noted the beta model literally generates faster than Mixtral 8x7B while costing drastically less. A second hands-on run measured 369–396 tok/s on longer generations. OrcaRouter's aggregate read of first-day testers is 280–500 tok/s sustained, with peaks above 500 under light concurrency. For calibration, Artificial Analysis measures Opus 5 at 57.1 tok/s at max effort and DeepSeek's own V4 Flash 0731 at ~108–140 tok/s.
- 3D and game generation got the loudest demos. Testers reported a complete Mario Kart-style racing game built inside DeepSeek's harness — generated sound effects, multiple maps — for roughly $1.10; a classical Chinese garden in Three.js complete with corridors and a pond with custom shaders in a single max-reasoning run; Minecraft-style voxel worlds; and spatial-reasoning tasks like dungeon navigation and exploded camera views. Geeky Gadgets summarised the sentiment as "a drastic leap" over V4 Flash's 3D output.
- Single-file front-end work is solid. One tester had it produce a fully responsive ocean-research dashboard in one HTML file, which matches what we saw from V4 Flash 0731 in our Gemini 3.8 Flash review-era testing of budget tiers: fast, correct, less decorative than Opus-class output.
- The complaints are specific. Reviewers flagged overthinking on straightforward prompts, and weaker performance on physics-based simulation — a rocket-launch test reportedly trailed GPT-6 Astra and Fable 5.1. OrcaRouter also noted early users grumbling that a model emitting tokens 2–3× faster "spends money faster" at an unchanged per-token price, and cautioned that per-task savings were unproven at launch because no third party had yet measured them.
- No independent score yet. At the time of writing, Artificial Analysis has not published an Intelligence Index result for V4.1 Flash. The nearest proxies from its own family are V4 Flash 0731 at 52 and V4 Pro 0813 at 53, against Opus 5's 63 — the highest score on the current index. Treat vendor gains as claims until that lands.
- Ecosystem support arrived day one. DeepSeek names WorkBuddy/CodeBuddy and OpenCode as launch partners fully supporting V4.1 Flash, and OpenRouter's traffic panel already shows it being driven by pi, DeepSeek's own Harness, Hermes Agent, Oh-My-Pi and Claude Code — with 151B prompt tokens logged.
Opus 5's field reports read differently: they are about judgement, not speed. Box ran it on its Complex Work Eval and measured 76% on due diligence versus Opus 4.8's 65%. Zapier reported Opus 5 hitting 100% on an end-to-end account-health churn-prevention sequence that prior models failed. Devin's team noted it approaches Fable-level performance on FrontierCode 1.1 "at half the cost," and Cursor reported it landing just under Fable 5 on CursorBench. Reviewers were less uniformly warm about its personality — Lenny's Newsletter titled its review "this model is brilliant (but annoying)", citing a "neurotic" streak in real coding sessions.
Pricing: 16.7× to 41.7×
| Cost component (per 1M tokens) | DeepSeek V4.1 Flash | Claude Opus 5 | Opus 5 premium |
|---|---|---|---|
| Input (cache miss) | $0.15 off-peak / $0.30 peak | $5.00 | 16.7× – 33.3× |
| Input (cache hit) | $0.003 off-peak / $0.006 peak | $0.50 | 83× – 167× |
| Output | $0.60 off-peak / $1.20 peak | $25.00 | 20.8× – 41.7× |
| Fast mode | — | ~2.5× speed, 2× price ($10 / $50) | — |
DeepSeek: peak hours are 01:00–04:00 and 06:00–10:00 UTC Monday–Friday; all other hours are off-peak at half price. New Flash rates took effect 04:00 UTC on September 10, 2026. Anthropic: $5/$25 with $0.50 cache reads, unchanged from Opus 4.8.
DeepSeek's September 10 cut is a real cut, not a reshuffle: on the Flash tier, off-peak cache-hit input dropped from $0.007 to $0.003, cache-miss input from $0.22 to $0.15, and output from $0.66 to $0.60 per million. It came 24 days after an August price increase that drew heavy developer criticism.
Sticker prices are only half the story, which is why cost per completed task matters. Artificial Analysis prices a full Intelligence Index run with Opus 5 at $2.03 per task at max effort, $1.23 at high, and $0.72 at medium — against roughly $0.11 per task for DeepSeek V4 Flash 0731 at the same index. V4.1 Flash has no measured figure yet, but the architecture's whole pitch is finishing the same work in fewer, faster steps.
What a Single Agent Task Costs
Take a mid-weight agent job of 1M input tokens and 250K output tokens. That is $11.25 on Opus 5, versus $0.60 on V4.1 Flash at peak and $0.30 off-peak — 18.75× and 37.5× cheaper. Push the numbers up and the gap is not linear, because Opus 5 bills 2,500% more per cached token: for a long agent loop where 90% of input is a cache hit, the cache line alone is $0.45 versus $0.0054 per million tokens.
The third-party anecdote that best captures the scale: a Hacker News commenter audited a TypeScript endpoint layer by layer (API, DTOs, service, database models) for $0.09 on a DeepSeek V4 model, where the same audit on Claude Opus 4.7 — same $5/$25 card as Opus 5 — was estimated at $9–13.
For planning your own workload, our AI model pricing calculator covers both cards, and the SWE-bench Pro leaderboard and MCP Atlas leaderboard track where each model sits on the coding and tool-orchestration axes.
Speed & Latency
DeepSeek figures are community-measured during the September 8–10 beta (World of AI peak 427 tok/s; separate hands-on 369–396 tok/s; aggregate first-day range 280–500 tok/s). Opus 5 figures are Artificial Analysis measurements at max, high and medium effort. Tokens per second is not an intrinsic model property — concurrency, effort setting and provider all move it.
Two caveats keep this chart honest. First, OpenRouter's independent monitoring puts V4.1 Flash at a P50 of 148 tok/s across all four providers, with DeepSeek's own endpoint averaging 137 tok/s and 1.37s P50 latency — so the viral 400+ numbers reflect best-case beta conditions, not typical production throughput. Second, Opus 5's 57.1 tok/s comes with an 85.93s time-to-first-token at max effort; at medium effort it starts answering in 6.70s and at high effort in 17.91s. For interactive work, the effort setting matters more than the decode rate.
Context, Output Length & Modalities
| Capability | DeepSeek V4.1 Flash | Claude Opus 5 |
|---|---|---|
| Context window | 1,000,000 tokens | 1,000,000 tokens |
| Max output | 384,000 tokens | 128,000 tokens |
| Image input | Yes (native, joint pre-training) | Yes |
| PDF input | Via API formats and tooling | Native in Messages API |
| KV cache footprint | 1/4 HBM, 1/8 SSD vs prior Flash gen | Prompt caching at $0.50/M reads |
| Concurrency | 2,500 requests | Rate-limited by tier |
| Data retention | Per DeepSeek's API terms | No retention requirement for general access |
Both models cover a million tokens of context, so the practical differentiator is output ceiling: 384K versus 128K. If you generate whole repos, long specs, or multi-file migrations in one pass, that is a real structural advantage for V4.1 Flash. If you need the model to report vulnerabilities rather than exploit them, or to be governed under Anthropic's alignment and safety stack — Opus 5 scored 2.3 on Anthropic's automated behavioural audit, the lowest misalignment score of its recent models — Opus 5 is the only option in this pair.
Verdict: Which Model for Which Job
Most teams should not pick one. The defensible setup in September 2026 is V4.1 Flash as the default lane — it is 19–38× cheaper per task, 2.5–7× faster to decode, and inside 2 points of Opus 5 on five agentic benchmarks — with Opus 5 escalated for the hard reasoning, long-context document analysis, and judgment-heavy tasks where its double-digit margins show up. Route by task shape, not model loyalty: if the task is "run this loop until the tests pass," DeepSeek; if it is "tell me what this stack of documents actually means," Opus 5.
Two things to watch. First, the missing Artificial Analysis score for V4.1 Flash: if it lands near 55–58, the "budget tier" framing dissolves. Second, V4.1 Pro, which DeepSeek has said is coming, will run on the same new architecture family — and the Flash-tier model has already set the bar it has to clear. Compare V4.1 Flash's siblings in our Opus 5 vs GPT-5.6 Sol, GLM-5.3 vs Opus 5, and Sonnet 5 vs DeepSeek V4 Pro write-ups, and follow the live boards: Terminal-Bench 2.1, Terminal-Bench 3.0, Terminal-Bench 4.0, FrontierCode v1.1, and FrontierBench v0.1.
Sources
- DeepSeek — Introducing DeepSeek-V4.1-Flash (launch note)
- DeepSeek API — Change log (2026-09-10 release and benchmark table)
- DeepSeek API — Models & Pricing (peak/off-peak rates, limits)
- Hugging Face — DeepSeek-V4.1-Flash weights & technical report
- Anthropic — Introducing Claude Opus 5
- Artificial Analysis — DeepSeek V4 Flash 0731 vs Claude Opus 5
- Artificial Analysis — Intelligence Index leaderboard
- OpenRouter — DeepSeek V4.1 Flash (pricing, throughput, traffic)
- Geeky Gadgets — DeepSeek V4.1 Flash review and performance test (427 tok/s)
- OrcaRouter — V4.1 Flash vs V4 Flash: same rate card, re-trained model
- Kie.ai — What is DeepSeek V4.1 Flash? Beta status and measured speeds
- Hacker News — DeepSeek launching V4.1 Flash cheaper and more capable than V4 Pro
- World of AI — DeepSeek V4.1 Flash hands-on (games, Three.js, 3D)
- Box — Claude Opus 5 on real enterprise work (Complex Work Eval)
- Lenny's Newsletter — Claude Opus 5 review: brilliant but annoying
- How I AI — Verdict on Claude Opus 5 after a 7-model benchmark
- Thomas Wiegold — DeepSeek V4 review on real code (cost-per-audit comparison)
- Quartz — DeepSeek V4-Flash is the cheapest major AI model to run