Gemini 3.7 Flash pricing graphic showing half-price discount ending before Gemini 3.5 Pro releaseGoogle's Gemini 3.7 Flash just got half price, but that deal disappears the same day Gemini 3.5 Pro is still nowhere to be found.
Gemini 3.7 Flash: Half Price Now, Full Price in 2027 AI & Enterprise Tech

Gemini 3.7 Flash Is Half Price. Read the Footnote First.

Google DeepMind shipped Gemini 3.7 Flash on August 13, 2026, its fourth Flash-tier model in nine weeks, at half the price of its predecessor. That discount expires December 31, 2026. And the model everyone actually asked for at I/O in May, Gemini 3.5 Pro, still hasn’t shipped.

If you’re choosing a model for coding agents or budgeting inference spend into 2027, both of those facts matter more than the launch headline. Here’s what Google’s own numbers say, what independent testing confirms, and what the pricing footnote is quietly telling you.

What Actually Shipped on August 13

Gemini 3.7 Flash went generally available in the Gemini API, Google AI Studio, Vertex AI, and Antigravity, Google’s coding-agent platform, on August 13, 2026, according to Google’s own Gemini API release notes. It also now powers Gemini Spark, Google’s productivity agent.

That’s four Flash-tier releases since Gemini 3.5 Flash debuted at I/O in May: 3.5 Flash, then 3.5 Flash-Lite and 3.6 Flash together on July 21, then 3.7 Flash on August 13. Twenty-three days between the last two. Google says the speed comes from algorithmic improvements, not a bigger base model.

Google AI Studio product lead Logan Kilpatrick framed it as a fast, targeted push rather than a ground-up rebuild.

“A strong intelligence increase, delivered in roughly three weeks through algorithmic work across Google DeepMind teams, focused on making the model feel more usable for real work.” Logan Kilpatrick, Product Lead, Google AI Studio, Google DeepMind, August 2026 launch announcement

Google’s own positioning line calls it “our most intelligent workhorse model yet for coding and agents.” Positioning aside, the two numbers worth caring about are what it costs and what it can actually do, and the answers to both come with asterisks.

The Pricing Trap: Half Price Until January 1

Gemini 3.7 Flash launched at $0.75 per million input tokens and $3.75 per million output tokens. That’s exactly half of what Gemini 3.6 Flash charges. It’s a genuinely good rate, and it’s live right now.

The part most launch-day coverage skipped: This pricing runs through December 31, 2026 only. On January 1, 2027, Gemini 3.7 Flash reverts to $1.50 input / $7.50 output per million tokens, the identical permanent rate Gemini 3.6 Flash has charged since its own July launch. The “half price” headline is a five-month promotional window, not a durable cost advantage.

If your team is modeling 2027 inference spend on today’s rate card, that model is wrong by roughly 2x. Build your cost projections around $1.50/$7.50, not $0.75/$3.75, for anything shipping past year-end.

There’s a second cost change buried in the same release: Google removed the “minimal” thinking tier, the cheap setting older Flash models used for high-volume classification work. “Low” is now the floor, and thinking tokens bill at the output rate even though the API only returns a summary of that reasoning. If your pipeline leaned on minimal-tier Flash for bulk, low-stakes calls, re-benchmark it. The effective cost floor just moved up even as the headline price moved down.

Gartner analyst Will Sommer flagged exactly this pattern months before this launch, and it applies directly here.

“Chief Product Officers should not confuse the deflation of commodity tokens with the democratization of frontier reasoning.” Will Sommer, Senior Director Analyst, Gartner, LLM inference economics forecast, March 2026

Sommer’s point is sharper once you factor in agentic workflows, which is exactly what Google is tuning 3.7 Flash for. Agent loops can multiply token consumption 5 to 30 times per task compared to a single chat completion. Cheaper tokens don’t necessarily mean a cheaper bill when the model is calling itself in a loop.

What Google’s Own Benchmarks Really Show

The capability jump from 3.6 Flash to 3.7 Flash is real, per Google’s own evals methodology page. DeepSWE v1.1 climbed from 49.0% to 65.3%. AutomationBench nearly doubled, from 17.0% to 30.4%.

Independent testing backs up at least the speed claim. Artificial Analysis clocked 3.7 Flash at 340.1 tokens per second in its standardized benchmarking workload, the fastest model the firm currently measures. But read Google’s own head-to-head comparison table against GPT-5.6 Terra in full, not cherry-picked, and the frontier-leadership framing softens fast.

MetricResult
Head-to-head rows won vs. GPT-5.6 Terra7 of 13
Rows lostTerminal-bench 2.1/3.0, OSWorld-2.0, DeepSWE v1.1, GDPval-AA v2
Price vs. GPT-5.6 TerraRoughly one-third the cost, per token

Notice which four rows it loses: the hardest agentic and terminal-use benchmarks, the exact category Google is marketing this model for. It’s a real value trade-off, not a clean win, and it’s Google’s own chart saying so.

There’s also a small but telling inconsistency worth flagging. Google’s 3.6 Flash model card lists its own DeepSWE v1.1 score as 48.6%. The 3.7 Flash launch blog rounds that same baseline to 49.0%. Minor on its own, but it’s a reminder that even single-vendor self-reported numbers are worth cross-checking against the vendor’s other documents, not just against competitors.

METR, the group that runs independent AI capability evaluations, has warned about a broader version of this problem.

“Benchmarks run without live human interaction can cause models to fail at tasks they could complete with minimal human guidance, making benchmarks unreliable proxies for real capability.” METR, Experienced Developer Study, July 2025

Translation for anyone building on this: treat DeepSWE and AutomationBench jumps as lab signals worth investigating, not as production-readiness guarantees. Run your own workload against it before you migrate.

The Elephant in the Room: Gemini 3.5 Pro

None of the Flash-tier sprint makes sense without the model that isn’t here. At I/O in May, Sundar Pichai told developers to give Google “until next month” for Gemini 3.5 Pro, implying a June release. It didn’t happen. As of this article’s publication, it still hasn’t.

Bloomberg reported on July 16, citing ten current and former Google employees, that 3.5 Pro was running months behind schedule, largely over coding-capability shortfalls. A late-June training-data update meant to fix that reportedly made results worse, not better. Later reporting sourced to the same chain indicates the problems ran deeper than a bad update: DeepMind concluded the original 3.5 Pro base model had structural failures in recursive tool-calling and SVG generation, scrapped it, and restarted pretraining from a native Gemini 3 foundation. That same reporting says DeepMind has already begun pretraining an entirely new flagship, Gemini 4, mentioned almost in passing in the July 21 announcement.

Pichai himself gave the first public crack in the story, back in May.

“A bit behind on agentic coding.” Sundar Pichai, CEO, Alphabet/Google, remarks at Google I/O, May 2026

Kilpatrick’s current line on 3.5 Pro is that the team is “testing with partners” and hopes to “land it soon.” That “soon” has now stretched past a second informal window with no date attached.

Wall Street has already priced in the uncertainty. Alphabet shares fell roughly 4.4% the day the Bloomberg delay report landed, an estimated $200 billion in market cap, on top of an earlier ~$225 billion drop in June tied to senior DeepMind researchers leaving for Anthropic and OpenAI. Combined, that’s close to $425 billion in Alphabet market value lost since late June with no change to reported revenue or earnings. Alphabet’s Q1 2026 results were strong (Google Cloud revenue up 63% year over year to $20 billion), which makes the point sharper: this is a narrative problem right now, not yet a fundamentals problem.

Our read: shipping four Flash models in nine weeks while the flagship reasoning tier stalls out looks less like a coincidence and more like a deliberate holding pattern, cover the volume segment on cost and speed while the harder model gets rebuilt underneath it. Google hasn’t confirmed that as strategy though, and it’s worth treating that framing as the most defensible inference from public facts, not as a confirmed internal decision. It could just as easily be ordinary engineering triage under deadline pressure.

What EU and UK Teams Need to Know

Buried in the model card, not the launch announcement, is a jurisdictional exclusion: Gemini 3.7 Flash is not available on the only consumer-facing surface it runs on in the EEA, UK, Switzerland, and Nigeria.

The timing isn’t nothing. The European Commission’s enforcement powers over general-purpose AI providers under the EU AI Act activated on August 2, 2026, penalties up to €15 million or 3% of global annual turnover, whichever is greater. Eleven days later, Google’s newest consumer AI model quietly excludes those exact jurisdictions from that surface. Google hasn’t stated a causal link publicly, but if you’re evaluating this model for an EU-facing product, plan around the exclusion now rather than discovering it in deployment.

The Bottom Line for Engineering Teams

If you’re on Gemini 3.6 Flash today, this is a real upgrade at a genuinely good price, for now. Three things to actually do with that:

  • Model your 2027 costs at $1.50/$7.50, not $0.75/$3.75. The current rate expires December 31, 2026.
  • Re-benchmark anything that ran on the old “minimal” thinking tier. It’s gone, and “low” now bills thinking tokens at the output rate.
  • Don’t lock a roadmap to a Gemini Pro milestone right now. Google has shipped zero Pro-tier models since Gemini 3 Pro in November 2025, despite promising 3.5 Pro for June 2026.

The broader question is whether this efficiency pivot holds. If Gemini 4, reportedly already in early pretraining, also slips, Google will have gone potentially 18 months or more between flagship releases while Anthropic and OpenAI keep a faster cadence. That’s a gap that compounds on reputation even if Flash-tier usage and revenue stay healthy in the meantime. Our related coverage on why enterprise AI inference costs aren’t actually falling and on Google’s agentic AI enterprise adoption gap both dig further into the pieces of this story we didn’t have room for here.


FAQ: Gemini 3.7 Flash and Gemini 3.5 Pro

Is Gemini 3.7 Flash actually cheaper than Gemini 3.6 Flash?

Only through December 31, 2026. Gemini 3.7 Flash launched at $0.75/$3.75 per million tokens, half of 3.6 Flash’s rate, but on January 1, 2027 it reverts to $1.50/$7.50, the same permanent rate 3.6 Flash has charged since July 2026.

When is Gemini 3.5 Pro coming out?

No confirmed date. Google promised it for June 2026 at I/O, but Bloomberg reported in July that coding-performance issues forced a delay, and later reporting indicates Google scrapped the original base model and restarted pretraining. As of August 15, 2026, it remains unreleased.

Is Gemini 3.7 Flash better than GPT-5.6 Terra for coding?

It’s close, not a clear win. On Google’s own 13-row comparison table, 3.7 Flash wins 7 rows but loses on the hardest agentic and terminal-use benchmarks to GPT-5.6 Terra, which costs over three times as much per token.

Why is Gemini 3.7 Flash not available in the EU or UK?

Google’s model card excludes the EEA, UK, Switzerland, and Nigeria from the consumer surface the model runs on. The exclusion lands 11 days after the EU AI Act’s enforcement powers over general-purpose AI providers activated on August 2, 2026.

What happened to Gemini Flash’s “minimal” thinking mode?

Gemini 3.7 Flash removed the “minimal” thinking tier used for cheap, high-volume classification tasks. “Low” is now the cheapest tier, and thinking tokens bill at the output rate even though only a summary is returned, raising the effective cost floor for simple workloads.


Gemini 3.7 Flash is a real, well-priced upgrade for teams already on the Flash tier, for the next four and a half months. What it isn’t is a replacement for the flagship model Google promised in May and still hasn’t shipped. Watch three things over the next 6 to 18 months: whether Gemini 3.5 Pro actually lands, whether the January price reset changes adoption at all, and whether Gemini 4’s pretraining run stays on schedule.

Want the next update the moment Gemini 3.5 Pro ships, or when the pricing resets in January? Subscribe to The Neural Loop at neuralwired.com/newsletter.

Leave a Reply

Your email address will not be published. Required fields are marked *