Google dropped Gemini 3.7 Flash this afternoon, its new workhorse model for coding and agents, and the most impressive thing is that only three weeks have passed since 3.6 Flash. Logan Kilpatrick, AI Studio's product lead, credits the jump to 'algorithmic improvements from teams across Google DeepMind' achieved in that window of time. The announcement, video included, went out at 7:07 PM CET:
It beats rivals costing twice as much at code
The benchmarks paint a small model fighting at the grown-ups' table: first place in Code Arena with a 1,588 Elo, ahead of Claude Sonnet 5 (1,541) and GPT-5.6 Terra (1,523), and the lead in FrontierCode too with 43.6%. In enterprise agent tasks the gap widens: 30.4% on AutomationBench against GPT-5.6 Terra's 23.6%.
It's not a clean sweep, on DeepSWE, the hardest software engineering test, GPT-5.6 Terra still rules with 69.6% against Google's 65.3%. And on the overall Artificial Analysis index the picture is a technical tie, 56 points to its two big rivals' 57. The read is that it isn't the smartest model in the world, it's the model that comes out smartest for what it costs.
The Flash family does the dirty work of the Gemini empire, the cheap model powering the quick answers of a platform that, as Google confirmed this very week, just passed 1 billion monthly users. Every cent Google shaves here gets multiplied by a scale no rival has.

$0.75 per million tokens, with rivals at double
Here's the real blow, $0.75 per million input tokens and $3.75 per million output, an introductory price that's half what 3.6 Flash debuted at three weeks ago. GPT-5.6 Terra charges $2 and $12 for the same, Muse Spark 1.2, $1.25 and $4.25. If tokens and their bookkeeping feel distant, the translation is simple, processing an entire book with Google's new model costs less than a coffee.
The model is live now in the API, AI Studio and the rest of Google's developer surfaces. And the trend it confirms is the trend of all of 2026, fast, cheap models are improving at a pace the flagships can no longer ignore, we covered how they work inside just recently. Three weeks between models. OpenAI's and Anthropic's answer is no longer measured in quarters, it's measured in days.
The conversation starts here
Sign in with a supporter account to comment. Sign in



Nobody has commented yet. Want to go first?