Gemini 3.7 Flash costs $0.75 per million input tokens today and $1.50 on January 1, 2027. The previous generation is not a cheaper place to hide, because it carries the same increase on the same date.
Google's pricing page states the paid-tier change plainly. Input is "$0.75 through December 31, 2026. $1.50 starting January 1, 2027." Output is "$3.75 through December 31, 2026. $7.50 starting January 1, 2027." The model went generally available on August 13, 2026 at that introductory rate.
Doubling is easy to read and easy to underestimate, because the interesting number is not the rate. It is what the rate does to the gap between models.
The premium for Flash more than triples
Gemini 3.5 Flash-Lite is listed at $0.30 input and $2.50 output with no date-conditional wording at all. So the two models drift apart on January 1 without Flash-Lite changing.
| Cost of 1M in plus 1M out | Through Dec 31, 2026 | From Jan 1, 2027 |
|---|---|---|
| 3.7 Flash | $4.50 | $9.00 |
| 3.5 Flash-Lite | $2.80 | $2.80 |
| Flash premium | +60.7% | +221.4% |
On real traffic, the model choice goes from a $7 decision to a $22 one
Take 10M input and 2M output tokens a month, a plausible shape for a product doing retrieval with short answers:
| Model | Through Dec 31, 2026 | From Jan 1, 2027 |
|---|---|---|
| 3.7 Flash | $15.00 | $30.00 |
| 3.5 Flash-Lite | $8.00 | $8.00 |
| Difference | $7.00 | $22.00 |
The choice between the two models is worth $7 a month today and $22 a month in January, on identical traffic. Anyone forecasting 2027 from a Q4 2026 invoice is understating the Flash line by half.
Staying on the older model saves nothing
The usual response to a price rise is to stay on the older model. That does not work here. gemini-3.6-flash lists the same $0.75 and $3.75 rates and the same January 1 increase to $1.50 and $7.50.
So the only documented route down the price ladder is Flash-Lite, and that is a different capability class rather than an earlier version of the same one. Which makes this a quality decision disguised as a billing decision: the cheaper option is cheaper because it is a smaller model, and whether it holds up on your workload is something only your evaluation can answer.
Evidence and limits
List prices for the paid standard tier, read on August 20, 2026 and re-checked against Google's pricing page on the morning this published, by a script in the repository that fails the release if the quoted figures have moved. Batch pricing, context caching discounts, free-tier limits, taxes, and negotiated terms are all excluded, and any of them can move a real bill. The token volumes above were chosen to be legible, not measured from a production system, so treat the dollar figures as ratios applied to your own numbers. No quality comparison between Flash and Flash-Lite was run here, which is exactly the gap the next section tells you to close yourself.
Four months is enough to answer this properly
There is time to answer this by measurement rather than by reacting in January. Capture a representative sample of your production prompts now, run it against both models, and record where Flash-Lite's answers are actually worse rather than assuming they are.
If they hold up, the migration is cheap while the price gap is still small. If they do not, you at least know what the extra $22 a month buys, which is a better position than discovering the doubled invoice first.
The same arithmetic makes context caching worth revisiting, because a discount on input tokens is worth twice as much once the rate behind it doubles. Whatever you decide, decide it before December 31, while the cheap version of the decision is still available.
