Skip to content
Nilay Kabariya

Entry No.046·AI··4 min read

Fix: Groq API error 413, 429 and model_not_found

Groq's 413 and 429 need different fixes. Its old Llama models now just say not found. All three reproduced on the free tier, with the fix for each.

by Nilay#groq#llm-api#nodeAI

The error, verbatim

Jump to the fix ↓
413 Request too large for model `openai/gpt-oss-20b` in organization `org_…` service tier `on_demand` on tokens per minute (TPM): Limit 8000, Requested 30074, please reduce your message size and try again.

429 Rate limit reached for model `openai/gpt-oss-20b` in organization `org_…` service tier `on_demand` on tokens per minute (TPM): Limit 8000, Used 4753, Requested 4824. Please try again in 11.8275s.

404 The model `llama-3.3-70b-versatile` does not exist or you do not have access to it.
✓ Reproduced on Windows 11 Pro

Tested on

groq-sdk
1.6.0
Tier
Groq free (on_demand)
Node
24.14.0
OS
Windows 11 Pro
Contents
  1. Which one do you have?01
  2. 1. 413: one request bigger than your limit02
  3. 2. 429: too much this minute03
  4. 3. 404 model_not_found: the Llama models are gone04
  5. What didn’t work05
  6. How this was tested06

Groq’s errors look alike, but they mean three different things. A 413 is one request that’s too big to ever fit your limit, so waiting won’t help. A 429 is several requests that together went over a limit, so waiting is exactly the fix. A 404 model_not_found usually means the model you copied from a tutorial has been shut down. Each one below was produced on a real free-tier Groq account.

Which one do you have?

Status Message starts with What it means Fix
413 Request too large for model One request is bigger than your per-minute token limit Send less per request (section 1)
429 Rate limit reached … tokens per minute (TPM) Your requests this minute add up to too many tokens Wait and retry (section 2)
429 Rate limit reached … requests per minute (RPM) Too many requests this minute Wait and retry (section 2)
404 The model … does not exist or you do not have access to it The model was shut down, or never existed Change the model name (section 3)
400 The model … has been decommissioned An older model that was shut down Change the model name (section 3)

1. 413: one request bigger than your limit

I sent openai/gpt-oss-20b about 30,000 tokens in one message:

413 Request too large for model `openai/gpt-oss-20b` in organization `org_…` service tier `on_demand`
on tokens per minute (TPM): Limit 8000, Requested 30074, please reduce your message size and try again.

The numbers are the whole story: the free tier allows 8,000 tokens per minute, and that one request asked for 30,074. It can never fit, however long you wait.

Make each request smaller than your TPM limit: shorter input, or split a long document into chunks and send them over several minutes.

Two things that don’t help on the free tier:

  • Switching models. openai/gpt-oss-20b, openai/gpt-oss-120b and qwen/qwen3.8-27b all reported the same 8000 token limit in their x-ratelimit-limit-tokens headers.
  • Lowering max_tokens. Groq didn’t count it toward this check: a 3,500-token prompt with max_tokens set to 1, 4,000 or 6,000 went through every time.

The message itself names the other option: a paid tier has higher limits.

2. 429: too much this minute

Two flavours, both reproduced.

Tokens per minute. Two requests of about 4,800 tokens each, back to back. The first succeeded; the second didn’t fit in the remaining 3,176:

429 Rate limit reached for model `openai/gpt-oss-20b` in organization `org_…` service tier `on_demand`
on tokens per minute (TPM): Limit 8000, Used 4753, Requested 4824. Please try again in 11.8275s.

Requests per minute. Forty tiny requests at once: 30 succeeded, 10 failed:

429 Rate limit reached for model `openai/gpt-oss-20b` in organization `org_…` service tier `on_demand`
on requests per minute (RPM): Limit 30, Used 30, Requested 1. Please try again in 2s.

Here waiting is the fix, and the message tells you how long (retry-after: 12 and retry-after: 2 in the headers).

The official SDK already retries 429s, but only twice. Give it more:

const groq = new Groq({ apiKey: process.env.GROQ_API_KEY, maxRetries: 6 });

For steady traffic, also keep your request rate under the limit instead of relying on retries. The free tier here allowed 30 requests per minute and 1,000 per day per model (x-ratelimit-limit-requests: 1000, resetting over the day). The daily cap isn’t something retries can fix.

3. 404 model_not_found: the Llama models are gone

The two Llama models most Groq examples use no longer work:

404 The model `llama-3.3-70b-versatile` does not exist or you do not have access to it.
404 The model `llama-3.1-8b-instant` does not exist or you do not have access to it.

Groq’s deprecations page says both were shut down on August 16, 2026. The error doesn’t say so. It reads like a typo or a permissions problem. Models retired earlier get a clearer message, with status 400:

400 The model `llama3-70b-8192` has been decommissioned and is no longer supported.

I got that for llama3-70b-8192, mixtral-8x7b-32768, gemma2-9b-it and deepseek-r1-distill-llama-70b.

Ask Groq which models your key can use, instead of trusting a list from a tutorial:

curl https://api.groq.com/openai/v1/models -H "Authorization: Bearer $GROQ_API_KEY"

Then use one from that list. These answered on my account: openai/gpt-oss-20b, openai/gpt-oss-120b and qwen/qwen3.8-27b.

What didn’t work

How this was tested

A free-tier Groq account (service tier on_demand), groq-sdk 1.6.0 and plain fetch on Node 24.14.0, Windows 11. Each error was produced deliberately with small, spaced-out requests: one 30,000-token message for the 413, two 4,800-token messages for the TPM 429, 40 parallel one-word messages for the RPM 429, and each retired model name for the 404 and 400 responses. Limits were read from Groq’s x-ratelimit-* headers. Error text is copied from the responses with the organization ID removed. Limits on paid tiers are higher and weren’t tested.

— N.K., end of entry No.046

Useful? Pass it on:Post on XFollow @EmotionalMatter

Related entries

  1. No.051

    Gemini 2.5 Flash "no longer available to new users" (404)

    Code that worked yesterday 404s with a new API key. Every old Gemini model name tested: three different messages, and why the model list won't warn you.

    > {

    AI3 min
  2. No.050

    Gemini API 429 RESOURCE_EXHAUSTED: why your limit is 0

    Gemini's 429 hides three different limits, and one of them never resets. Each hit on purpose with a free key, plus the 503 high demand error.

    > {

    AI4 min
  3. No.045

    Fix: APIConnectionError "Connection error." (OpenAI, Anthropic)

    OpenAI's and Anthropic's SDKs both say only Connection error. Five real causes reproduced, from a local model that isn't running to antivirus HTTPS scanning.

    > APIConnectionError: Connection error.

    AI5 min

Post card · Newsletter

Get the next fix in your inbox.

One short email when a new entry is published. No spam, never shared, and you can leave any time.

— Nilay

or follow by RSSor on X

By subscribing you agree to the privacy note. One click to leave.

tip: paste the exact error text