/v1 routes return the OpenAI-style error envelope below. /v1/messages and /v1/messages/count_tokens return the Anthropic error envelope instead.
type is the broad family, code is the specific reason, and param names the request field when one caused it. Branch on type and code, never on message text. Some errors add fields directly inside error, such as paused_until, limit, or credit amounts, never nested under a details property.
For a 401, the wire code is invalid_api_key. Clients may still tolerate auth_error as a legacy spelling for that status.
On OpenAI-compatible routes, type is one of exactly six values.
Failures inside the gateway return the route’s documented envelope, including a failure we did not anticipate. On OpenAI-compatible routes, that arrives as
500 internal_error rather than an empty body. The platform edge can answer before the gateway with no JSON envelope: see Payload too large and HTML 403 security checkpoint.
OpenAI-compatible error responses carry X-Request-Id, matching request_id. Anthropic-compatible responses carry both request-id and x-request-id. Quote a request id when you contact support.
Should I retry?
The absence of the retry headers is a deliberate signal, not an oversight. A JSON400, 401, 402, 403, 404 or 413 needs a different request or different credentials, so none carries timing. An HTML 403 needs a pause before retrying.
Rotation cannot leave more than 20 live customer API keys in a scope, twice the normal cap of 10. A rotation that would exceed this ceiling returns HTTP 400 with top-level code api_key_rotation_overlap_limit, without changing existing keys. Revoke an unneeded live key in that scope, then retry. While another key operation for the same scope is still in progress, creating or rotating a key can answer HTTP 503 with top-level code api_key_mint_busy and a Retry-After: 2 header; nothing changed, wait two seconds and send the same request again.
Status and code reference
Every
503 a hosted model can return is in the table below.
The 503 codes
A503 means the gateway could not safely serve the request. During a rate-limit store outage, each server applies rate limits on its own instead of returning limiter_unavailable, but hosted admission and Idempotency-Key checks can still return a retryable 503. All of them carry Retry-After except upstream_http_error, which forwards one only when the model supplied it, and hosted_inference_unavailable, which carries none. For datastore_unavailable, hosted_evidence_unavailable, hosted_admission_unavailable and the connection gate’s limiter_unavailable, the wait is randomized between 1 and 4 seconds, up to 8 under heavy load.
Coding plan limit
A serving coding plan returns402 plan_limit_reached when a plan window is at its limit and neither credits nor Standby can serve the request. The response carries x-should-retry: false and no Retry-After. It is a billing refusal; temporary Standby capacity refusals use 429 with retry timing.
The fields sit directly inside error, not under details. This illustrative Pro response was constructed with the plan’s exact window limits. Credits after limits are disabled and today’s Standby is exhausted:
credits.blocker is disabled, balance_negative, key_limit, empty, cap, or period_unavailable. An owner can receive Keep full speed on credits, Add funds, Raise the key limit, Raise the cap, or Contact support as the next step. A member receives Ask an owner.
The funding headers report x-runinfra-funding: refused. /v1/messages rewrites the body to error.type: "billing_error" and the message text, with request_id; the other plan fields are absent. In Messages clients such as Claude Code, read the message: it carries the limit, reset, Standby condition, and next step without an error.code. See Anthropic Messages.
Envelopes that carry more than the table
402 insufficient_credits, with the balance and the top-up link
402 insufficient_credits, with the balance and the top-up link
Credit fields are in cents. These values are an example.Add at least
required_cents - current_balance_cents, then retry. topup_url is absolute so it resolves from a terminal, an SDK exception, or a log line, none of which have an origin to resolve a relative path against.When the workspace has a coding plan that is not paying for requests (a payment is overdue, or a payment dispute paused it), message ends with the plan’s state and its next step, such as paying the overdue amount, or contacting support.A free_window_ended_at field appears only when the model you called was on a promotional free window that closed within the last seven days, as a UTC RFC 3339 instant describing the same close the message states in words. When there is no recently closed window the field is absent rather than null, which is the usual case. The same field accompanies a 402 account_frozen under the same rule. A model’s price, and whether it is inside a free window, live on that model’s page, never here.402 spend_limit_reached, with the limit, the spend, and the reset instant
402 spend_limit_reached, with the limit, the spend, and the reset instant
A per-key monthly spending limit is optional and set by the key’s owner under Settings, API keys. It is enforced on the key’s settled spend in the current calendar month (UTC), so a request already in flight when the limit is reached may still complete. Amounts are in cents. These values are an example.This describes the pay-as-you-go key limit. On a serving coding plan, the key limit blocks credits after plan limits; eligible plan or Standby usage can still run. See key payment settings. Other keys in the workspace are unaffected. Raise or remove the limit on the key, use another key, or retry after
resets_at.503 hosted_model_paused, with the next availability check
503 hosted_model_paused, with the next availability check
Retry-After is in seconds: the time until the next availability check, capped at one hour. Once that check time has passed, it is 60 seconds on every endpoint, so a client polls about once a minute instead of every second. The status and code stay the same.A paused model is not a missing model. It still appears in GET /v1/models with availability: "paused", so treat this as retry and poll, never as a reason to drop the model id from your configuration.400 hosted_parameter_not_supported on response_format
400 hosted_parameter_not_supported on response_format
reasoning_effort: "none" beside response_format, or drop the format. The rule is per model, and the affected models are named in what each model supports. Nothing is charged for the refusal, which is the point: the alternative is a 200 carrying a reply your parser rejects, billed in full.429 hosted_admission_congested, a capacity shed
429 hosted_admission_congested, a capacity shed
Under a momentary burst the gateway waits only briefly for capacity and then sheds the request rather than let it wait longer, so it never completes. That is a temporary capacity condition, not an outage and not a quota, so it answers This answer applies on every
429./v1 route, so the same condition never surfaces as a different status elsewhere. The message varies with where the shed happened; the status, type and code do not. Unlike a hosted admission limit, a shed carries no limit field and no X-Hosted-Limit-Scope header, because there is no quota to size against.401 or 403 on a key that looks correct
401 or 403 on a key that looks correct
Browser or device sign-in with
runinfra login creates a terminal key. It has the same rp_ prefix and 40 character body as a workspace API key and it cannot call the inference API, so it answers authentication_error with the message “This API key is internal and cannot be used for customer inference.” Create a workspace API key in Settings, API keys.HTML 403 security checkpoint
An HTML403 security checkpoint instead of a JSON error means automatic traffic protection flagged your network. The response carries x-vercel-mitigated: challenge. Wait a few minutes before retrying, lower the request burst, and contact support if it persists.
Payload too large
The 3.5MB ceiling is enforced on the declaredContent-Length and again while the body streams, so an understated Content-Length does not bypass it. It is a limit on the encoded request and is separate from the model context window: split the prompt across requests or send less context.
Well above the limit, around 4MB and beyond, the platform edge rejects the request before the gateway sees it, with a plain-text FUNCTION_PAYLOAD_TOO_LARGE response carrying no JSON envelope and no request id. Treat it as the same instruction: shrink the body.
Retry rules
- Keep the same
Idempotency-Keywhen you retry, streaming or not. - Respect
Retry-After, orRetry-After-Msif your client reads it, on every response that carries them. - On the hosted admission
429codes the value is an estimate, capped at 8 seconds, and it changes between refusals. Add jitter if many workers share one key. - Change the request or the credentials before retrying a JSON
400,401,402,403,404or413. They carry no timing, and that is the signal. - Do not replace the key after
422 idempotency_replay_unavailable. - A streaming retry never replays the delivered tokens. The same key returns
409while the original is open, then the terminal usage and cost once it settles. See Idempotent retries.
Related
Rate limits
What each 429 means, and how to back off.
Idempotent retries
Retry safely, streaming or not.
Troubleshooting
Symptom first, from the response you actually got.