Skip to main content
Four independent layers can refuse a Model APIs request, and they run in this order:
  1. A connection gate, before your credential is read.
  2. A per-key request window, requests per minute for that key.
  3. Hosted admission, concurrency and workspace tokens per minute on a shared model.
  4. The model at capacity, after admission, when every server the request reached is full.
All four answer HTTP 429 with the standard error envelope. The connection gate and the per-key window both use rate_limit_exceeded, hosted admission uses its own codes, and a model at capacity uses hosted_model_at_capacity. /v1/messages uses the Anthropic envelope, but the same Retry-After headers and limits apply.

Read the header, never a constant

Every 429 carries Retry-After. The per-key window, hosted admission and the model at capacity also send Retry-After-Ms.
Retry-After-Ms is the same instant in milliseconds. Read it first if your client supports it, and the OpenAI and Anthropic SDKs both do: it is the only one of the two that can express a wait shorter than a second without rounding up. Retry-After is the RFC 9110 header in whole seconds, always at least 1, always rounded up, so it never advises you to retry early.

Layer 1, the connection gate

Before the API reads your key, a gate bounds how fast unrecognized credentials can be looked up, so a flood of invalid keys cannot crowd out real traffic. It applies per source address and, while the rate-limit store is healthy, across the service as a whole. There is no warm-up and nothing to request: a new key works on its first call. What is bounded is the rate at which credentials that have never authenticated can be presented. With a healthy store, once a key succeeds it stops counting against that budget, and steady traffic from a working key never meets this gate. You are most likely to see it when starting many workers at once with a brand new key, or when a script is retrying a key that is simply wrong. With a healthy store, rolling a new key out on one request before fanning out avoids it entirely. The budgets are not published, because this is a defensive control and it is tuned. A gate refusal carries no X-RateLimit-* headers, which sets it apart from the per-key window. If the rate-limit store stops answering, each API server applies the same budgets in memory, without sharing counts with other servers. A key the server already trusted keeps bypassing the gate on that server for up to 15 minutes; other keys get that server’s new-key budgets. After three consecutive store failures within 10 seconds, a server’s rate limits stop calling the store for 5 seconds and then probe it. A successful probe restores shared counting on that server. When the store is not configured or a request cannot be evaluated, the gate returns 503 limiter_unavailable rather than admitting unbounded work. Its Retry-After is randomized between 1 and 4 seconds, up to 8 seconds under heavy load, so refused clients do not all return at once.

Layer 2, the per-key request window

A rolling 60 second window while the rate-limit store is healthy. A key with no custom limit follows the workspace’s current default, so a workspace upgrade applies without rotating the key. A custom per-key limit stays in place, clamped to the workspace maximum.
The window slides, so capacity returns gradually as individual requests age past 60 seconds. There is no clock edge where the whole budget refills at once, which is why pacing evenly beats bursting. During a store outage, each server counts the key’s requests in its own fixed 60 second window, and the X-RateLimit-* headers describe that server’s count. Going over the limit there returns 429. If the store is not configured, the request returns 503 limiter_unavailable rather than being served unmetered. Hosted admission and requests with an Idempotency-Key still need the store, so they can return a retryable 503 during an outage. See the 503 codes.

Layer 3, hosted admission

Hosted admissionchecked in this orderModel capacityConfigured per modelDeepSeek V4 Flash: 81 API key8 -> 2 in flightmax(1, floor(capacity / 4))2 Workspace8 -> 2 in flightsame value as the key gate3 Large-prompt shednot a size limitvery large prompts only, near full capacity4 Shared model8 in flightthe full capacity, all workspaces5 Workspace tokens per minute200,000 tokensper-model budget, rolling 60s window429with Retry-After, X-Hosted-Limit-Scope, and X-Hosted-Limit-Value naming the gateThe token reservation is an upper bound (one token per UTF-8 byte) and is reconciledto the served token counts after the response.
Hosted admissionchecked in this orderModel capacityConfigured per modelDeepSeek V4 Flash: 81 API key8 -> 2 in flightmax(1, floor(capacity / 4))2 Workspace8 -> 2 in flightsame value as the key gate3 Large-prompt shednot a size limitvery large prompts only, near full capacity4 Shared model8 in flightthe full capacity, all workspaces5 Workspace tokens per minute200,000 tokensper-model budget, rolling 60s window429with Retry-After, X-Hosted-Limit-Scope, and X-Hosted-Limit-Value naming the gateThe token reservation is an upper bound (one token per UTF-8 byte) and is reconciledto the served token counts after the response.
The gates run in the order of the table below. Each one counts per model. A hosted refusal carries limit inside error, plus headers naming the gate that stopped you.
Every hosted refusal uses the same envelope. This is the API-key concurrency case, and the numbers in it are an example:
Always read error.limit and X-Hosted-Limit-Value from the response you got, never a number from this page. They change with capacity. Hosted admission 429 responses also carry X-RateLimit-Limit-Tokens, X-RateLimit-Remaining-Tokens, and X-RateLimit-Reset-Tokens when the limiter has a measured workspace token snapshot. This includes concurrency refusals, so a client can back off with both the refusing concurrency limit and the remaining token budget in view.

Layer 4, the model at capacity

A request can pass every gate above and still find the model’s servers full. When the server a request reaches is at capacity, the gateway tries one other server if the model has more than one. If every server it reached was full, the request answers 429 rate_limit_error with code hosted_model_at_capacity and a Retry-After header. Wait for Retry-After and retry unchanged; the same request succeeds once a slot frees. It is not a limit on your workspace, so it carries no limit field and no X-Hosted-Limit-Scope header, and nothing is charged. An SDK with retries enabled handles it like any other 429.

Handling a 429 in code

An SDK with retries enabled does the right thing with no code from you, because every layer sends Retry-After. If you back off yourself, add jitter when many workers share a key so they do not all return at the same instant.

Three cases worth planning for

  • A 429 before any request of yours has succeeded, on a key you just created, is the connection gate rather than a quota. Retry once after Retry-After with a single request, then resume normal concurrency.
  • One request that alone exceeds the TPM limit is the case where waiting cannot help. Its refusal still carries Retry-After, but retrying it unchanged fails again. Reduce the input or the maximum output tokens before retrying.
  • Adding API keys does not buy capacity. Every key in a workspace draws on one shared concurrency allowance, so a second key does not raise it. Read the current allowance for a model from max_concurrent_requests_per_api_key in GET /v1/models rather than assuming a fixed figure: it tracks that model’s capacity, which changes as a model moves onto larger capacity. The allowance is what each workspace is held to once the model is busy, which keeps a busy model shared rather than taken by whoever arrived first. While the model has capacity to spare, requests beyond it are admitted up to a higher ceiling, and the allowance applies again as the model fills; a refusal always quotes the limit that applied at that moment in error.limit. Hosted token-rate admission uses the same request token estimate as billing, not a byte count.

Coding plans and Standby

A coding plan’s 5-hour and weekly limits meter usage separately from request and capacity limits. Exhausting plan funding returns a 402 plan limit, with x-should-retry: false and no Retry-After. Requests still running reserve their share of a limit until they finish. When those reservations alone hold the rest of a 5-hour or weekly limit, the credits cap, or today’s Standby, a new request waits on the server for up to 30 seconds. If it is still held, the request gets 429 plan_window_busy with Retry-After and x-should-retry: true. error.details.window is 5h, weekly, credits_cap, or standby_day. The request is not charged and never moves to credits or Standby. While a plan serves covered requests, Pro and Team have a temporary floor of at least 20 million tokens per minute per model. Net plan payments contribute to the earned token tier. The temporary floor ends when the plan stops serving. Plan-funded requests have their own concurrency allowance per model, published as max_concurrent_plan_requests_per_workspace in GET /v1/models. Standby serves chat only and depends on model capacity. A temporary plan or Standby refusal uses these 429 codes and includes retry timing: Read the applied value from error.limit or X-Hosted-Limit-Value. A daily Standby amount does not guarantee that the model can admit a request now.

Errors

Every status, code, and whether a retry can help.

Models

Discover the limits that bind your key at boot.