Gemini 3.7 Flash Launches, But the Real Story Is Reliability
Google shipped Gemini 3.7 Flash on August 13, 2026, just 23 days after 3.6 Flash underperformed rivals, positioning it as the fix. The coding gains are real: DeepSWE v1.1 jumped from 49.0% to 65.3%, and FrontierCode 1.1 Main climbed to 43.6%, edging out Claude Sonnet 5’s 42.7%. Introductory pricing is $0.75 per million input tokens and $3.75 per million output, half of 3.6 Flash’s rate.

The Catch Nobody Flagged on Launch Day

Two details are missing from Google’s own announcement thread. First, that $0.75/$3.75 price doubles to $1.50/$7.50 on January 1, 2027, the exact rate 3.6 Flash launched at. It is a four-and-a-half-month promotion, not a new baseline. Second, Google’s footnote excludes the European Economic Area, the United Kingdom, Switzerland, and Nigeria from Gemini Spark, the only consumer surface running 3.7 Flash today. Developers reach the model directly through the API, AI Studio, and Google’s code execution tools in Gemini Notebook regardless of region, but everyday Workspace users in those excluded markets are stuck on the older model inside Spark.

Our Take

The benchmark story and the reliability story are not the same story, and Google is letting the first one drown out the second. A model can post genuine gains on DeepSWE and FrontierCode while still failing at tasks that require zero reasoning: creating a Google Doc element inside a Google Doc, converting a basic LaTeX file, building a calendar event from an email’s plain content. Those are not edge cases requiring frontier intelligence. They are the exact “everyday Workspace tool use” Google’s own Spark pitch promises, and they are failing on the older model that most users in Google’s own ecosystem are still running, including the newly excluded regions.

Benchmarks measure what a model can do under ideal conditions with the right harness. They say nothing about what happens when the same model is wired into Gmail and asked to do the boring thing correctly, the thing it is actually shipped to do. Google’s pattern this year, three model releases in roughly two months, plus expanding agent creation to all mobile users and pushing into robotics with Gemini Robotics 2.0, reads as a company racing to claim benchmark leadership across every surface simultaneously. That race is happening while its original Gemini technical co-leads have both left for OpenAI within months of each other, and Sergey Brin has reportedly stepped back into active development. Shipping fast is not the problem. Shipping fast while the product still lies about completing basic tasks, and calling that intelligence, is.

Follow us on Bluesky, LinkedIn, X, and Telegram to Get Instant Updates