Gemini 3.7 Flash (Google DeepMind, August 13, 2026) just landed — three weeks after its predecessor — and it's the fastest reasoning model on the market. Grok 4.6 (SpaceXAI, August 12, 2026) released the day before it, with the opposite pitch: maximum intelligence per dollar. Two launches, 24 hours apart, two very different philosophies.

This guide compares the two across benchmarks, pricing, speed, context, and modalities. Figures are drawn from Google's launch materials, xAI's launch table, Artificial Analysis, and independent coverage.

TL;DR: Gemini 3.7 Flash is the speed & multimodal king: ~340 tokens/second (4× Grok's 85.8), $0.75/$3.75 introductory pricing, 1M context, and video/audio input. Grok 4.6 is the intelligence king: 61 vs 56 on the AA Intelligence Index and a crushing 1753 vs 1525 GDPVal-AA Elo lead. If you need throughput, video, and budget — Gemini. If you need reasoning depth and professional-grade knowledge work — Grok.

At a Glance

CategoryGrok 4.6Gemini 3.7 Flash
Release dateAugust 12, 2026August 13, 2026
DeveloperSpaceXAIGoogle DeepMind
Context window500,000 tokens1,048,576 tokens
Max output65,536 tokens
Input modalitiesText, imageText, image, video, audio, PDF
Input price$2 / 1M$0.75 / 1M (intro)
Output price$6 / 1M$3.75 / 1M (intro)
Post-intro price$1.50 / $7.50 (from Jan 1, 2027)
ReasoningLow, medium, high, xhighConfigurable (medium / high)
Model IDgrok-4.6gemini-3.7-flash

Benchmark Comparison

Grok 4.6 leads on overall intelligence and knowledge work; Gemini 3.7 Flash is strikingly close on long-horizon software engineering despite its "Flash" label.

EvaluationGrok 4.6Gemini 3.7 FlashLeader
AA Intelligence Index6156Grok 4.6
GDPVal-AA v2 (Elo)17531525Grok 4.6
Terminal-Bench 2.188.4%85.8%Grok 4.6
DeepSWE v1.165.9%65.3%≈ Tie
FrontierCode 1.161.3% (Ext)43.6% (Main)Different splits
τ³-Banking50.7%Grok 4.6
AutomationBench30.4%Gemini 3.7
GDP.pdf34.0%Gemini 3.7

Sources: Artificial Analysis (Aug 2026), Google launch chart, xAI launch table. Vendor/third-party reported; harnesses differ. FrontierCode splits (Extended vs Main) are not directly comparable.

Coding Benchmarks Head-to-Head

DeepSWE is the headline surprise: Gemini 3.7 Flash jumped from 49.0% (3.6 Flash) to 65.3% — statistically tied with Grok 4.6's 65.9%. Terminal-Bench 2.1 goes to Grok by 2.6 points. Google's Flash line is no longer just "the fast cheap one" — it's now genuinely competitive on long-horizon engineering.

Knowledge-Work Elo

Here the gap is decisive: Grok 4.6's 1753 GDPVal-AA Elo outclasses Gemini's 1525 by 228 points. For professional documents, structured analysis, and judgment-heavy knowledge work, Grok is in a different tier. Gemini's wins are elsewhere — AutomationBench (30.4%, ~3× Claude Sonnet 5) and GDP.pdf document comprehension (34.0%).

Capability Radar

Radar values are normalized (0–100) to the higher of the two models on each axis: Intelligence = AA Index (61), Knowledge Work = GDPVal-AA v2 Elo (1753), Terminal = Terminal-Bench 2.1, SWE = DeepSWE v1.1, Speed = output t/s, Value = cost per AA task (inverted).

Pricing: Gemini's Intro Discount Is the Real Story

Through December 31, 2026, Gemini 3.7 Flash costs $0.75/$3.75 per 1M tokens — 62% cheaper than Grok's $2/$6 on input and 37% cheaper on output. After that, it reverts to $1.50/$7.50. Cost per completed task tells the same story:

Cost componentGrok 4.6Gemini 3.7 Flash
Input (per 1M)$2$0.75 (intro)
Output (per 1M)$6$3.75 (intro)
Cost per AA Index task$0.84$0.40 (high) / $0.26 (medium)

Consider a hypothetical agent task using 1M input + 250K output tokens:

At current rates, that task costs roughly $3.50 on Grok 4.6 vs. ~$1.69 on Gemini 3.7 Flash — half the price. Gemini's cost per task ($0.40) is also less than half of Grok's ($0.84). But note the calendar: the Gemini discount expires January 1, 2027, and doubles. Model any 2027 cost model on the full $1.50/$7.50 rate.

Speed & Latency

This is Gemini's home turf. It's the fastest reasoning model Artificial Analysis has measured:

Metric (Artificial Analysis)Grok 4.6 (high)Gemini 3.7 Flash
Output speed85.8 t/s~340 t/s (high) / ~274 t/s (med)
Time to first token32.3 s— (not published)
Time per AA Index task~1.7 minutes

Gemini 3.7 Flash generates ~4× faster than Grok 4.6 (340 vs 85.8 t/s) — nearly 3× the speed of GPT-5.6 Terra and GLM-5.2, placing it on the Intelligence vs. Time-per-Task Pareto frontier. If your agent loops wait on decode time, this single number may dominate every other consideration. (Google hasn't published a time-to-first-token figure; Grok's is 32.3s.)

Context & Modalities: Gemini's Other Two Advantages

Beyond speed, Gemini 3.7 Flash brings two structural advantages:

  1. Double the context: 1,048,576 tokens vs Grok's 500K — with a strong 97.0% on GDM-MRCR v2 long-context retrieval at 128K.
  2. True multimodality: Gemini accepts video and audio input alongside text, images, and PDFs. Grok 4.6 is text + image only. For video summarization, meeting analysis, or media-heavy pipelines, Gemini is the only option of the two.

Verdict: Which Should You Choose?

Choose Gemini 3.7 Flash if your workload is volume-driven, latency-sensitive, or multimodal: agent loops that wait on decode, video/audio processing, or high-throughput coding. At $0.40 per task and 340 t/s until year-end, it's the best speed-per-dollar in the market — just re-check the economics when the intro price doubles in January 2027.
Choose Grok 4.6 if you need reasoning depth and professional knowledge work: 5 more points on the Intelligence Index, a 228-point GDPVal-AA Elo lead, and stronger terminal-work performance. It costs more and runs slower, but it thinks harder — and unlike Gemini, its $2/$6 pricing has no expiry date.

The natural setup for many teams is complementary routing: Gemini 3.7 Flash for the high-volume, latency-sensitive, and multimodal lanes; Grok 4.6 for the hard reasoning and knowledge-work lanes. Different tools, different jobs — and together they cover more of the frontier for less than either of the premium models.

Sources