The error, verbatim
Jump to the fix ↓413 Request too large for model `openai/gpt-oss-20b` in organization `org_…` service tier `on_demand` on tokens per minute (TPM): Limit 8000, Requested 30074, please reduce your message size and try again.
429 Rate limit reached for model `openai/gpt-oss-20b` in organization `org_…` service tier `on_demand` on tokens per minute (TPM): Limit 8000, Used 4753, Requested 4824. Please try again in 11.8275s.
404 The model `llama-3.3-70b-versatile` does not exist or you do not have access to it.
Tested on
- groq-sdk
- 1.6.0
- Tier
- Groq free (on_demand)
- Node
- 24.14.0
- OS
- Windows 11 Pro
Contents
Groq’s errors look alike, but they mean three different things. A 413 is one request that’s too big to ever fit your limit, so waiting won’t help. A 429 is several requests that together went over a limit, so waiting is exactly the fix. A 404 model_not_found usually means the model you copied from a tutorial has been shut down. Each one below was produced on a real free-tier Groq account.
Which one do you have?
| Status | Message starts with | What it means | Fix |
|---|---|---|---|
| 413 | Request too large for model |
One request is bigger than your per-minute token limit | Send less per request (section 1) |
| 429 | Rate limit reached … tokens per minute (TPM) |
Your requests this minute add up to too many tokens | Wait and retry (section 2) |
| 429 | Rate limit reached … requests per minute (RPM) |
Too many requests this minute | Wait and retry (section 2) |
| 404 | The model … does not exist or you do not have access to it |
The model was shut down, or never existed | Change the model name (section 3) |
| 400 | The model … has been decommissioned |
An older model that was shut down | Change the model name (section 3) |
1. 413: one request bigger than your limit
I sent openai/gpt-oss-20b about 30,000 tokens in one message:
413 Request too large for model `openai/gpt-oss-20b` in organization `org_…` service tier `on_demand`
on tokens per minute (TPM): Limit 8000, Requested 30074, please reduce your message size and try again.
The numbers are the whole story: the free tier allows 8,000 tokens per minute, and that one request asked for 30,074. It can never fit, however long you wait.
Make each request smaller than your TPM limit: shorter input, or split a long document into chunks and send them over several minutes.
Two things that don’t help on the free tier:
- Switching models.
openai/gpt-oss-20b,openai/gpt-oss-120bandqwen/qwen3.8-27ball reported the same8000token limit in theirx-ratelimit-limit-tokensheaders. - Lowering
max_tokens. Groq didn’t count it toward this check: a 3,500-token prompt withmax_tokensset to 1, 4,000 or 6,000 went through every time.
The message itself names the other option: a paid tier has higher limits.
2. 429: too much this minute
Two flavours, both reproduced.
Tokens per minute. Two requests of about 4,800 tokens each, back to back. The first succeeded; the second didn’t fit in the remaining 3,176:
429 Rate limit reached for model `openai/gpt-oss-20b` in organization `org_…` service tier `on_demand`
on tokens per minute (TPM): Limit 8000, Used 4753, Requested 4824. Please try again in 11.8275s.
Requests per minute. Forty tiny requests at once: 30 succeeded, 10 failed:
429 Rate limit reached for model `openai/gpt-oss-20b` in organization `org_…` service tier `on_demand`
on requests per minute (RPM): Limit 30, Used 30, Requested 1. Please try again in 2s.
Here waiting is the fix, and the message tells you how long (retry-after: 12 and retry-after: 2 in the headers).
The official SDK already retries 429s, but only twice. Give it more:
const groq = new Groq({ apiKey: process.env.GROQ_API_KEY, maxRetries: 6 });For steady traffic, also keep your request rate under the limit instead of relying on retries. The free tier here allowed 30 requests per minute and 1,000 per day per model (x-ratelimit-limit-requests: 1000, resetting over the day). The daily cap isn’t something retries can fix.
3. 404 model_not_found: the Llama models are gone
The two Llama models most Groq examples use no longer work:
404 The model `llama-3.3-70b-versatile` does not exist or you do not have access to it.
404 The model `llama-3.1-8b-instant` does not exist or you do not have access to it.
Groq’s deprecations page says both were shut down on August 16, 2026. The error doesn’t say so. It reads like a typo or a permissions problem. Models retired earlier get a clearer message, with status 400:
400 The model `llama3-70b-8192` has been decommissioned and is no longer supported.
I got that for llama3-70b-8192, mixtral-8x7b-32768, gemma2-9b-it and deepseek-r1-distill-llama-70b.
Ask Groq which models your key can use, instead of trusting a list from a tutorial:
curl https://api.groq.com/openai/v1/models -H "Authorization: Bearer $GROQ_API_KEY"Then use one from that list. These answered on my account: openai/gpt-oss-20b, openai/gpt-oss-120b and qwen/qwen3.8-27b.
What didn’t work
How this was tested
A free-tier Groq account (service tier on_demand), groq-sdk 1.6.0 and plain fetch on Node 24.14.0, Windows 11. Each error was produced deliberately with small, spaced-out requests: one 30,000-token message for the 413, two 4,800-token messages for the TPM 429, 40 parallel one-word messages for the RPM 429, and each retired model name for the 404 and 400 responses. Limits were read from Groq’s x-ratelimit-* headers. Error text is copied from the responses with the organization ID removed. Limits on paid tiers are higher and weren’t tested.
— N.K., end of entry No.046