The Gains Are Concentrated in Coding and Agents
By Google's own published figures, FrontierCode 1.1 Main goes from 34.4% to 43.6%, DeepSWE v1.1 from 49.0% to 65.3%, AutomationBench from 17.0% to 30.4%, and Arena.ai's WebDev Elo from 1538 to 1588. Google attributes the gains to improvements in multi-step planning and tool use, which points at getting code right on the first pass and at agent runs that actually finish. The model is generally available through the Gemini API and reachable from AI Studio and Antigravity.
The Discount Is Real, but Check the Expiry
Introductory pricing is $0.75 per million input tokens, $3.75 per million output, and $0.075 per million for context caching — half what 3.6 Flash cost at launch. That rate runs only through December 31, 2026, after which it doubles to $1.50 input and $7.50 output. For reference, Google's own comparison table lists Claude Sonnet 5 at $2.00 / $10.00 and GPT-5.6 Terra at $2.00 / $12.00. So the price advantage in this capability tier is substantial until year-end and narrows sharply afterward — any team budgeting long-term against this rate should write the doubling into the model.
Don't Read Only the Headline Numbers
The same official data shows where it trails: Gemini 3.7 Flash scores 85.8% on Terminal-bench 2.1 against 87.4% for GPT-5.6 Terra, which also leads on Terminal-bench 3.0 and OSWorld-2.0, while Claude Sonnet 5 wins Agent's Last Exam 33.3% to 26.3%. One report also notes that three of the five headline improvements come from evaluations run by AI companies rather than independent parties. The fair reading is that this update makes the mid price tier considerably more competitive on coding and agent work, not that it displaces the more expensive models outright. Google also gave no release date for Gemini 3.5 Pro, with 3.7 Flash shipping ahead of the Pro tier.
via: Google's announcement, "Gemini 3.7 Flash: our most intelligent workhorse model", The Decoder report, VentureBeat report; verified 2026-08-15