Back to the blog
AI & Pricing

Google Is Cutting Free Gemini Down to Flash-Lite on October 9. Your App's Free Tier Was Never a Contract Either

October 7, 2026 6 min read

On October 6, The Verge reported that Google is about to take Flash and Pro away from people who use Gemini for free. Google's own help center confirms it, in a page titled "Changes to Gemini model access and limits": starting October 9, people who use the Gemini app on a personal account without an AI subscription will only have Flash-Lite.

The rest of the ladder moves too. Google AI Plus, $4.99 a month in the US, keeps Flash-Lite and Flash but loses Pro, on a date Google says each subscriber will receive by email. Pro stays with AI Pro, at $19.99 a month, and AI Ultra, from $99.99. On the day I write this, Google's US subscription page still describes the free plan as "Access to 3.6 Flash" and "Varying access to 3.1 Pro." From October 9, the smallest of the three models becomes the whole free offer.

I build products on Claude, not Gemini. I still read it twice, because a developer's first question is obvious: does this hit the API?

What is changing, and what is not

The answer is no, at least not this week. Google scopes the change to "Gemini Apps when you use a personal account." I checked the developer pages. The Gemini API pricing page, last updated October 6, still lists Gemini 3.8 Flash as "Free of charge" on the free tier, input and output, and the same for Gemini 3.5 Flash-Lite. The API release notes, updated the same day, announce no free tier change.

So if your side project calls the Gemini API with a free key, October 9 is a normal day. If your users talk to Gemini through Google's app, they lose Flash.

On limits, Google's public pages give no figures on either side. In the app, usage for users over 18 has been "compute-based" since May 17: limits refresh every five hours until you reach a weekly cap, AI Plus gets twice the standard limit and AI Pro four times. For the API, the rate limits page prints no free tier figures at all. It sends you to AI Studio to see the limits of each project, and adds: "Specified rate limits are not guaranteed and actual capacity may vary."

The API free tier moves too, just more quietly

The app change got a headline. The API changes mostly got a line.

The current Pro model, Gemini 3.1 Pro Preview, is already marked "Not available" on the API free tier. On September 18, the release notes announced that access to the 2.5 models is now limited "to users who have actively used them in the past," and told new projects to use 3.5 Flash-Lite or 3.8 Flash. In January, on Google's developer forum, Logan Kilpatrick replied for Google to people whose daily quota had dropped in the AI Studio playground: "We did reduce the limits for free in the UI." He added: "I expect the limits to continue to go down over time." That was the playground, not the API, but the direction was stated plainly.

Even the paid price carries a date. Gemini 3.8 Flash costs $0.75 per million input tokens and $3.75 per million output tokens "through December 31, 2026." From January 1, 2027, the same page lists $1.50 and $7.50. The doubling is already published, less than three months ahead.

None of this is a scandal. A free tier is a vendor's spare capacity, packaged as a product, and it shrinks when the spare capacity does. The pricing page says so in its last line: "Rate limits are subject to change."

A dependency you did not sign for

Pay for an API and you get terms, a price list and, sometimes, notice. Build on a free tier and you get none of that, and you rarely write down that you depend on it. It never shows up on a bill, so it never shows up in a planning meeting.

I start most side projects on free or cheap tiers too. The failure I find in my own older code is not the free tier itself. It is that the model ID sits as a string literal in several files, the rate-limit error lands in a generic catch block, and nobody knows how many tokens each feature burns. When the tier changes, a pricing decision turns into a debugging session.

There is also a quieter line. Google's pricing page says content sent on the free tier is "used to improve our products." On the paid tier, it is not. For the LegalTech I build, where client documents describe inventions nobody has filed yet, that line alone rules a free tier out.

What I build so a vanishing tier is a config change

Four things, all cheap to add early.

One door for every model call. Every call goes through a single function that takes a feature name, not a model name. A config file maps each feature to a provider and a model. When a model leaves a free tier, I change one line of config. If a search for a model ID finds it in more than one file, that is the first thing to fix.

A fallback that has actually run. Each feature gets a second model, ideally from another provider, and a written rule for what happens when the first one says no: a 429 RESOURCE_EXHAUSTED (the error Google documents for its spend limits), a timeout, a model that no longer answers. Decide per feature whether to fail open or closed. The small judge model I run in a side project fails open for a while, then starts rejecting after eight consecutive failures, because what it lets through cannot be taken back. Then push some real traffic through the fallback every week. A fallback that has never served a request is a guess.

Tokens per feature, logged on every call. Store input and output tokens next to the feature name. That table turns "the free tier is gone" into a number in five minutes. Say a feature reads 20 million input tokens and writes 2 million output tokens a month on Gemini 3.8 Flash. At the paid price through December 31, that is 20 × $0.75 + 2 × $3.75 = $22.50 a month. From January 1, it is 20 × $1.50 + 2 × $7.50 = $45. Without the table, you are guessing at your own bill.

A cap and an alert before the first paid request. Since March, AI Studio lets you set a monthly spend cap per project. Google labels it experimental and warns that billing data can lag by around ten minutes, so overages remain possible. Above that, each billing tier has its own monthly ceiling, $250 on Tier 1. And moving to paid now means prepaying at least $5 of credit: the billing docs say the API "will only serve requests if you have a positive Prepay credit balance." An empty balance is an outage too, so alert on the balance, not only on the spend.

Google just showed the pattern

The funny part is that Google's change is exactly this architecture, applied to its own users. When compute is expensive, send the traffic that pays least to the cheapest model, keep the better models for the traffic that pays, and change the mapping with a table, not a rewrite. Flash-Lite becomes the floor.

Your app deserves the same knob. Decide now which features could run on a Lite-class model if they had to, and test them on one, once. If the answer is "all of them, slightly worse," a free tier disappearing is an ordinary Tuesday. If the answer is "none of them," you have found a dependency worth paying for, before someone else decides the date.

On Monday, search your code for model IDs. The number of hits is the number of places one pricing email can break.

Sources

A project like this one?

I design and deploy products like this. Let's talk.

Let's talk