Skip to main content
Most /v1 routes return the OpenAI-style error envelope below. /v1/messages and /v1/messages/count_tokens return the Anthropic error envelope instead.
type is the broad family, code is the specific reason, and param names the request field when one caused it. Branch on type and code, never on message text. Some errors add fields directly inside error, such as paused_until, limit, or credit amounts, never nested under a details property. For a 401, the wire code is invalid_api_key. Clients may still tolerate auth_error as a legacy spelling for that status. On OpenAI-compatible routes, type is one of exactly six values. Failures inside the gateway return the route’s documented envelope, including a failure we did not anticipate. On OpenAI-compatible routes, that arrives as 500 internal_error rather than an empty body. The platform edge can answer before the gateway with no JSON envelope: see Payload too large and HTML 403 security checkpoint. OpenAI-compatible error responses carry X-Request-Id, matching request_id. Anthropic-compatible responses carry both request-id and x-request-id. Quote a request id when you contact support.

Should I retry?

The absence of the retry headers is a deliberate signal, not an oversight. A JSON 400, 401, 402, 403, 404 or 413 needs a different request or different credentials, so none carries timing. An HTML 403 needs a pause before retrying. Rotation cannot leave more than 20 live customer API keys in a scope, twice the normal cap of 10. A rotation that would exceed this ceiling returns HTTP 400 with top-level code api_key_rotation_overlap_limit, without changing existing keys. Revoke an unneeded live key in that scope, then retry. While another key operation for the same scope is still in progress, creating or rotating a key can answer HTTP 503 with top-level code api_key_mint_busy and a Retry-After: 2 header; nothing changed, wait two seconds and send the same request again.

Status and code reference

Every 503 a hosted model can return is in the table below.

The 503 codes

A 503 means the gateway could not safely serve the request. During a rate-limit store outage, each server applies rate limits on its own instead of returning limiter_unavailable, but hosted admission and Idempotency-Key checks can still return a retryable 503. All of them carry Retry-After except upstream_http_error, which forwards one only when the model supplied it, and hosted_inference_unavailable, which carries none. For datastore_unavailable, hosted_evidence_unavailable, hosted_admission_unavailable and the connection gate’s limiter_unavailable, the wait is randomized between 1 and 4 seconds, up to 8 under heavy load.

Coding plan limit

A serving coding plan returns 402 plan_limit_reached when a plan window is at its limit and neither credits nor Standby can serve the request. The response carries x-should-retry: false and no Retry-After. It is a billing refusal; temporary Standby capacity refusals use 429 with retry timing. The fields sit directly inside error, not under details. This illustrative Pro response was constructed with the plan’s exact window limits. Credits after limits are disabled and today’s Standby is exhausted:
credits.blocker is disabled, balance_negative, key_limit, empty, cap, or period_unavailable. An owner can receive Keep full speed on credits, Add funds, Raise the key limit, Raise the cap, or Contact support as the next step. A member receives Ask an owner. The funding headers report x-runinfra-funding: refused. /v1/messages rewrites the body to error.type: "billing_error" and the message text, with request_id; the other plan fields are absent. In Messages clients such as Claude Code, read the message: it carries the limit, reset, Standby condition, and next step without an error.code. See Anthropic Messages.

Envelopes that carry more than the table

A per-key monthly spending limit is optional and set by the key’s owner under Settings, API keys. It is enforced on the key’s settled spend in the current calendar month (UTC), so a request already in flight when the limit is reached may still complete. Amounts are in cents. These values are an example.
This describes the pay-as-you-go key limit. On a serving coding plan, the key limit blocks credits after plan limits; eligible plan or Standby usage can still run. See key payment settings. Other keys in the workspace are unaffected. Raise or remove the limit on the key, use another key, or retry after resets_at.
Retry-After is in seconds: the time until the next availability check, capped at one hour. Once that check time has passed, it is 60 seconds on every endpoint, so a client polls about once a minute instead of every second. The status and code stay the same.A paused model is not a missing model. It still appears in GET /v1/models with availability: "paused", so treat this as retry and poll, never as a reason to drop the model id from your configuration.
A new key here would start new inference work and a second charge. Contact support with request_id instead.
Send the request again with reasoning_effort: "none" beside response_format, or drop the format. The rule is per model, and the affected models are named in what each model supports. Nothing is charged for the refusal, which is the point: the alternative is a 200 carrying a reply your parser rejects, billed in full.
Under a momentary burst the gateway waits only briefly for capacity and then sheds the request rather than let it wait longer, so it never completes. That is a temporary capacity condition, not an outage and not a quota, so it answers 429.
This answer applies on every /v1 route, so the same condition never surfaces as a different status elsewhere. The message varies with where the shed happened; the status, type and code do not. Unlike a hosted admission limit, a shed carries no limit field and no X-Hosted-Limit-Scope header, because there is no quota to size against.
Browser or device sign-in with runinfra login creates a terminal key. It has the same rp_ prefix and 40 character body as a workspace API key and it cannot call the inference API, so it answers authentication_error with the message “This API key is internal and cannot be used for customer inference.” Create a workspace API key in Settings, API keys.

HTML 403 security checkpoint

An HTML 403 security checkpoint instead of a JSON error means automatic traffic protection flagged your network. The response carries x-vercel-mitigated: challenge. Wait a few minutes before retrying, lower the request burst, and contact support if it persists.

Payload too large

The 3.5MB ceiling is enforced on the declared Content-Length and again while the body streams, so an understated Content-Length does not bypass it. It is a limit on the encoded request and is separate from the model context window: split the prompt across requests or send less context. Well above the limit, around 4MB and beyond, the platform edge rejects the request before the gateway sees it, with a plain-text FUNCTION_PAYLOAD_TOO_LARGE response carrying no JSON envelope and no request id. Treat it as the same instruction: shrink the body.

Retry rules

  • Keep the same Idempotency-Key when you retry, streaming or not.
  • Respect Retry-After, or Retry-After-Ms if your client reads it, on every response that carries them.
  • On the hosted admission 429 codes the value is an estimate, capped at 8 seconds, and it changes between refusals. Add jitter if many workers share one key.
  • Change the request or the credentials before retrying a JSON 400, 401, 402, 403, 404 or 413. They carry no timing, and that is the signal.
  • Do not replace the key after 422 idempotency_replay_unavailable.
  • A streaming retry never replays the delivered tokens. The same key returns 409 while the original is open, then the terminal usage and cost once it settles. See Idempotent retries.

Rate limits

What each 429 means, and how to back off.

Idempotent retries

Retry safely, streaming or not.

Troubleshooting

Symptom first, from the response you actually got.