July 2026 has been a blockbuster month for frontier AI. On June 30, Anthropic dropped Claude Sonnet 5 — the "most agentic Sonnet yet" — and three weeks later, on July 21, Google DeepMind answered with Gemini 3.6 Flash, a token-efficient workhorse that delivers near-Pro quality at Flash pricing. Both models represent the new mid-tier: powerful enough for production, cheap enough to run at scale.
But which one actually wins? We dug into every official benchmark, every system card, and every independent evaluation to give you the clearest picture possible. Every number in this post is sourced — from Google DeepMind's model card, Anthropic's Sonnet 5 System Card, and independent leaderboards like BenchLM and Artificial Analysis. No estimates, no guesses.
📋 Table of Contents
🔬 Head-to-Head: 8 Shared Benchmarks
Google DeepMind's own Gemini 3.6 Flash product page includes a comparison table that pits 3.6 Flash directly against Claude Sonnet 5 (among others). These are the only benchmarks where both models were evaluated on the same harness, making them the fairest comparison available.
| Benchmark | 🔵 Gemini 3.6 Flash | 🟠 Claude Sonnet 5 | Winner |
|---|---|---|---|
| SWE-Bench Pro | 58.7% | 63.2% | 🟠 Claude +4.5 |
| DeepSWE v1.1 | 49.0% | 54.0% | 🟠 Claude +5.0 |
| Terminal-Bench 2.1 | 78.0% | 80.4% | 🟠 Claude +2.4 |
| MLE-Bench | 63.9% | 66.9% | 🟠 Claude +3.0 |
| GDPval-AA v2 (Elo) | 1,421 | 1,618 | 🟠 Claude +197 |
| OSWorld-Verified | 83.0% | 81.2% | 🔵 Gemini +1.8 |
| CharXiv Reasoning | 85.2% | 77.0% | 🔵 Gemini +8.2 |
| GDM-MRCR v2 (128k) | 91.8% | 71.6% | 🔵 Gemini +20.2 |
Figure 1: Head-to-head comparison across 7 percentage-based benchmarks. Bold outlines indicate the winner on each metric. GDPval-AA v2 shown separately below due to its different scale (Elo).
Score: Claude Sonnet 5 leads 5–3 on shared benchmarks. But the split reveals something more interesting than a simple win count. Claude dominates coding and knowledge work. Gemini dominates multimodal reasoning and long-context retrieval — and by massive margins.
🎯 Capability Radar: Where Each Model Shines
The radar chart below normalizes all shared benchmarks to a 0–100% scale, revealing the shape of each model's intelligence. Gemini 3.6 Flash is spiky — extraordinary at visual reasoning and long context, competitive at coding. Claude Sonnet 5 is rounder — strong across the board, with particular depth in knowledge work and agentic coding.
Figure 2: Radar chart showing the capability profile of each model. Gemini dominates the right hemisphere (multimodal + long context); Claude dominates the left (coding + knowledge).
💻 Coding & Agentic Performance
This is Claude Sonnet 5's home turf. Anthropic explicitly positioned it as "the most agentic Sonnet yet," and the numbers back it up.
Claude Sonnet 5 — Coding Benchmarks
| Benchmark | Sonnet 5 | Sonnet 4.6 | Gain |
|---|---|---|---|
| SWE-bench Verified | 85.2% | 79.6% | +5.6 |
| SWE-bench Pro | 63.2% | 58.1% | +5.1 |
| Terminal-Bench 2.1 | 80.4% | 67.0% | +13.4 ⚡ |
| FrontierCode v1 | 38.8% | 15.1% | +23.7 🔥 |
| CursorBench | 57.0% | 49.0% | +8.0 |
| DeepSWE v1.1 | 54.0% | — | — |
The Terminal-Bench 2.1 result is the headline: Sonnet 5 at 80.4% actually beats Opus 4.8 (74.6%) on the same harness. This is the first time a Sonnet-class model has surpassed the concurrent Opus flagship on any benchmark. And on FrontierCode v1, Sonnet 5 more than doubles Sonnet 4.6's score — from 15.1% to 38.8%.
Gemini 3.6 Flash isn't standing still either. Its DeepSWE score jumped from 37% (3.5 Flash) to 49% — a 12-point gain. MLE-Bench rose from 49.7% to 63.9%, a 14.2-point leap that actually puts it within striking distance of Claude Sonnet 5's 66.9%.
Figure 3: Claude Sonnet 5's additional benchmarks from the Anthropic System Card (June 30, 2026).
🧠 Knowledge Work & Reasoning
On GDPval-AA v2 — the Elo-style benchmark that measures economically valuable knowledge work — Claude Sonnet 5 posts a staggering 1,618. That's not just ahead of Gemini 3.6 Flash (1,421). It actually beats Opus 4.8 (1,615), Anthropic's $5/M token flagship. This is the first time a mid-tier model has outscored the concurrent flagship on knowledge work.
Figure: GDPval-AA v2 — Knowledge Work Elo. Claude Sonnet 5 not only beats Gemini 3.6 Flash by 197 points, but also surpasses Anthropic's own Opus 4.8 flagship (1,615).
On Humanity's Last Exam (HLE), the hardest academic reasoning benchmark in circulation, Sonnet 5 scores 57.4% with tools — essentially tied with Opus 4.8 (57.9%) and a massive 10.6-point jump over Sonnet 4.6. Without tools, it scores 43.2%.
Gemini 3.6 Flash doesn't publish HLE scores directly, but its AA Intelligence Index sits at 50 — roughly flat with 3.5 Flash, confirming Google positioned this as an efficiency release rather than a raw intelligence leap.
👁️ Multimodal & Long Context: Gemini's Territory
This is where the comparison flips dramatically. Gemini 3.6 Flash is natively multimodal — it accepts text, images, audio, video, and PDFs as input. Claude Sonnet 5 is text-plus-image only.
On CharXiv Reasoning (information synthesis from complex charts), Gemini 3.6 Flash scores 85.2% without tools and 89.4% with tools. Claude Sonnet 5 manages 77.0% — an 8.2-point gap. This matters enormously for financial analysis, scientific research, and any workflow involving charts and diagrams.
But the real blowout is long-context retrieval. On GDM-MRCR v2 at 128k tokens, Gemini scores 91.8% versus Claude's 71.6% — a 20.2-point chasm. At the full 1 million token depth, Gemini scores 54.0%, roughly double what its predecessor managed. Claude Sonnet 5 has no published score at this depth.
On OSWorld-Verified (agentic computer use), Gemini edges ahead 83.0% to 81.2% — a narrow but real lead in the benchmark that tests real-world GUI navigation.
Figure 4: Gemini 3.6 Flash's additional benchmarks from Google DeepMind (July 21, 2026).
💰 Pricing & Token Economics
Figure 5: API pricing comparison. Gemini 3.6 Flash is 2× cheaper on input and 2× cheaper on output.
| Pricing Factor | 🔵 Gemini 3.6 Flash | 🟠 Claude Sonnet 5 |
|---|---|---|
| Input ($/1M tokens) | $1.50 | $3.00 ($2.00 intro) |
| Output ($/1M tokens) | $7.50 | $15.00 ($10.00 intro) |
| Cached Input ($/1M) | $0.15 | $0.30 ($0.20 intro) |
| Token Efficiency | ~17% fewer tokens vs 3.5 Flash | 1.0–1.35× more tokens (new tokenizer) |
| Effective Cost/Task | 💰 Lower | 💰💰 Higher |
Gemini 3.6 Flash is the clear cost winner. At $1.50/$7.50 per million tokens, it's half the price of Claude Sonnet 5 on both input and output. And with ~17% fewer output tokens per task (per the Artificial Analysis Index), the real-world cost gap is even wider than the sticker prices suggest.
Claude Sonnet 5's pricing has an additional wrinkle: a new tokenizer that maps the same input to 1.0–1.35× more tokens. Anthropic set introductory pricing ($2/$10 through August 31, 2026) to make the transition roughly cost-neutral for Sonnet 4.6 users. After that, standard rates ($3/$15) apply — meaning Sonnet 5 actually costs more per character than Sonnet 4.6 did.
📋 Specifications at a Glance
| Specification | 🔵 Gemini 3.6 Flash | 🟠 Claude Sonnet 5 |
|---|---|---|
| Release Date | July 21, 2026 | June 30, 2026 |
| Developer | Google DeepMind | Anthropic |
| Context Window | 1,048,576 tokens | 1,000,000 tokens |
| Max Output | 65,536 tokens | 64,000 tokens |
| Input Modalities | Text, Image, Audio, Video, PDF | Text, Image |
| Output | Text | Text |
| Knowledge Cutoff | March 2026 | Early 2026 |
| Reasoning | Configurable (thinking budget) | Extended thinking (effort dial) |
| Tool Use | Function calling, Search, Computer Use | Function calling, Computer Use, MCP |
| Arena Elo (Text) | 1,482 | ~1,510 (est.) |
| API Model ID | gemini-3.6-flash | claude-sonnet-5 |
🏆 The Verdict
There's no single winner here — the right model depends entirely on your workload. Here's our breakdown:
🔵 Choose Gemini 3.6 Flash if…
- Multimodal is critical — you process audio, video, or PDFs
- Long-context retrieval matters — 91.8% at 128k is unmatched
- Cost is a primary concern — 2× cheaper on both input and output
- Computer use / GUI agents — leads at 83.0% on OSWorld
- Chart-heavy analysis — 85.2% on CharXiv vs 77.0%
- You need the freshest knowledge — March 2026 cutoff
🟠 Choose Claude Sonnet 5 if…
- Agentic coding is your focus — 85.2% SWE-bench Verified, 80.4% Terminal-Bench
- Knowledge work quality is paramount — 1,618 GDPval-AA beats even Opus 4.8
- You need deep reasoning — 57.4% HLE with tools
- Brownfield debugging — traces failures to root causes
- Long-running autonomous agents — sustains focus over complex multi-step tasks
- You're in the Anthropic ecosystem — Claude Code, MCP, Bedrock
Bottom Line
Claude Sonnet 5 is the stronger model overall for coding and knowledge work — it leads on 5 of 8 shared benchmarks and posts genuinely shocking numbers on Terminal-Bench and GDPval-AA. The fact that a $3/M Sonnet-class model can beat a $5/M Opus-class model on multiple benchmarks is a structural shift in the AI pricing landscape.
Gemini 3.6 Flash is the smarter choice for multimodal and cost-sensitive workloads. It's half the price, natively multimodal, and absolutely dominates long-context retrieval. For teams building RAG pipelines, document analysis, or computer-use agents, Gemini 3.6 Flash delivers more capability per dollar than anything else on the market.
The real takeaway? The mid-tier is now the sweet spot. Both of these models deliver capability that would have required flagship pricing just six months ago. You no longer need to pay Opus or Pro prices to get production-grade AI. The question isn't which model is "best" — it's which model is best for your specific workload.
📚 Sources
- Google DeepMind — Gemini 3.6 Flash product page (benchmark table, July 2026)
- Gemini 3.6 Flash Model Card (Google DeepMind, July 2026)
- Introducing Claude Sonnet 5 (Anthropic, June 30, 2026)
- Claude Sonnet 5 System Card (Anthropic, June 30, 2026)
- Claude Sonnet 5 Benchmarks Explained (Vellum, June 30, 2026)
- Claude Sonnet 5 Benchmarks: Every Score Explained (Emergent, July 2026)
- BenchLM — Gemini 3.6 Flash (July 2026)
- BenchLM — SWE-bench Verified Leaderboard (July 2026)
- Google Released Gemini 3.6 Flash (DataNorth, July 22, 2026)
- Gemini 3.6 Flash Review (Analytics Vidhya, July 2026)
- Artificial Analysis — Gemini 3.6 Flash (July 2026)
- Claude Sonnet 5 vs Sonnet 4.6 (CodingFleet, June 2026)