The error, verbatim
Jump to the fix ↓{
"error": {
"code": 429,
"message": "You exceeded your current quota, please check your plan and billing details. ...
* Quota exceeded for metric: generativelanguage.googleapis.com/generate_content_free_tier_requests, limit: 0, model: gemini-3.1-pro
Please retry in 10h17m5.723541104s.",
"status": "RESOURCE_EXHAUSTED"
}
}
Tested on
- @google/genai
- 2.27.0
- API tier
- Free (AI Studio key)
- Node / OS
- 24.14.0 / Windows 11 Pro
- Date
- October 7, 2026
Contents
Every Gemini quota error looks the same at first: HTTP 429, "status": "RESOURCE_EXHAUSTED", “You exceeded your current quota”. The part that matters is one line further down, in the * Quota exceeded for metric lines: the limit number and the model. With a free AI Studio key I hit three different ones on purpose, and each one needs a different fix:
| The message says | What it means | Retry delay it gave me |
|---|---|---|
limit: 0, model: gemini-3.1-pro |
Your tier has no quota for this model | 10h17m (waiting doesn’t help) |
limit: 5, model: gemini-3.5-flash |
Requests per minute used up | 46s |
limit: 20, model: gemini-3.5-flash |
Requests per day used up | 9h54m, until 00:00 UTC |
The details array names the quota exactly. The per-minute one was GenerateRequestsPerMinutePerProjectPerModel-FreeTier, the daily one GenerateRequestsPerDayPerProjectPerModel-FreeTier.
limit: 0: the model isn’t in your tier
This was the first call I made with the key. It wasn’t a burst: a single request to gemini-3.1-pro-preview came back 429 straight away. Same for gemini-pro-latest. On the free tier these models have a quota of zero. The error still says “Please retry in 10h17m”, but that only marks when the daily counters reset. Retrying at that time gets the same limit: 0.
Use a model that your tier has quota for. With this free key, gemini-3.5-flash, gemini-3.5-flash-lite and gemini-flash-latest answered normally:
const res = await ai.models.generateContent({
model: 'gemini-3.5-flash',
contents: 'hi',
});If you need the Pro model specifically, that means a paid tier (billing turned on for the Google Cloud project behind the key). I didn’t test that: it costs money, and it’s the only way past a limit: 0.
Per-minute limit: wait the time it gives you
Sending one tiny request after another to gemini-3.5-flash, calls 1 to 7 succeeded, and call 8 failed 8.4 seconds in:
* Quota exceeded for metric: generativelanguage.googleapis.com/generate_content_free_tier_requests, limit: 5, model: gemini-3.5-flash
Please retry in 46.130728812s.
The limit says 5, but seven got through first, so don’t use the number to pace yourself exactly. The @google/genai SDK (2.27.0) threw this straight to my code: it didn’t retry on its own.
Wait the retryDelay from the error’s RetryInfo detail, then retry. But cap the wait, because a daily limit sends a delay of hours:
import { GoogleGenAI, ApiError } from '@google/genai';
const ai = new GoogleGenAI({ apiKey: process.env.GEMINI_API_KEY });
async function generateWithRetry(params, maxWaits = 4) {
for (let attempt = 0; ; attempt++) {
try {
return await ai.models.generateContent(params);
} catch (err) {
if (!(err instanceof ApiError) || attempt >= maxWaits) throw err;
if (err.status === 503) { // "high demand": back off and try again
await new Promise((r) => setTimeout(r, 2000 * 2 ** attempt));
continue;
}
if (err.status !== 429) throw err;
const body = JSON.parse(err.message).error;
if (/limit: 0\b/.test(body.message)) throw err; // no quota for this model at all
const delay = body.details?.find((d) => d['@type']?.endsWith('RetryInfo'))?.retryDelay ?? '10s';
const ms = Math.ceil(parseFloat(delay) * 1000) + 500;
if (ms > 120_000) throw err; // a daily quota: the delay runs to midnight UTC
await new Promise((r) => setTimeout(r, ms));
}
}
}Daily limit: it resets at midnight UTC, not Pacific time
My test calls used up gemini-3.5-flash’s free daily quota (limit: 20). The retry delay pointed at 2026-10-08 00:00:13 UTC. That’s 8 PM in New York and 5 PM in San Francisco, not midnight Pacific as a lot of older answers say.
Switch to a different model until the reset. Each model has its own counters: right after gemini-3.5-flash hit its daily limit, nine calls in a row to gemini-3.5-flash-lite all succeeded.
Not a 429: “This model is currently experiencing high demand”
While I was testing, gemini-3.8-flash (the newest Flash) returned this on every call, and later gemini-3.5-flash did too, for a few calls:
{
"error": {
"code": 503,
"message": "This model is currently experiencing high demand. Spikes in demand are usually temporary. Please try again later.",
"status": "UNAVAILABLE"
}
}
This is Google being overloaded, not your quota. There’s no retry delay in it, and other models kept working at the same moment. Calls that did go through during that period were slow, too: one took 56 seconds.
Back off and retry, or use another model for a while. The helper above waits 2, 4, then 8 seconds. In my run the 503s stopped after those three waits (the next answer was the daily-limit 429, because by then my test calls had used up the day’s quota).
What didn’t work
How this was tested
A free Google AI Studio API key on October 7, 2026, with @google/genai 2.27.0 on Node 24.14.0 (Windows 11). Every request was one word with maxOutputTokens: 1, sent one after another, never in parallel. Limits were reached on purpose: a Pro model for limit: 0, a fast loop for the per-minute limit, and the day’s test calls for the daily one. The limits shown are the free tier’s on that day. Google changes them, so read the numbers in your own error message. I didn’t test paid tiers, Vertex AI, or the Gemini CLI, whose 429s come from a different quota.
The code for every case above is public, so you can run it yourself: gemini-429 in nk-repro.
— N.K., end of entry No.050