> ## Documentation Index
> Fetch the complete documentation index at: https://runinfra.ai/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Errors

> The error envelope, every status and code, and whether retrying can help.

Most `/v1` routes return the OpenAI-style error envelope below. `/v1/messages` and `/v1/messages/count_tokens` return the [Anthropic error envelope](/docs/api-reference/anthropic-messages#errors) instead.

```json theme={"dark"}
{
  "error": {
    "message": "Invalid API key",
    "type": "authentication_error",
    "param": null,
    "code": "invalid_api_key"
  },
  "request_id": "req_example_0000"
}
```

`type` is the broad family, `code` is the specific reason, and `param` names the request field when one caused it. Branch on `type` and `code`, never on `message` text. Some errors add fields directly inside `error`, such as `paused_until`, `limit`, or credit amounts, never nested under a `details` property.

For a `401`, the wire code is `invalid_api_key`. Clients may still tolerate `auth_error` as a legacy spelling for that status.

On OpenAI-compatible routes, `type` is one of exactly six values.

| `error.type` | Statuses that use it |
| - | - |
| `invalid_request_error` | 400, 405, 409, 413, 422 |
| `authentication_error` | 400, 401, 403 |
| `permission_error` | 402, 403 |
| `not_found_error` | 404 |
| `rate_limit_error` | 429 |
| `api_error` | 500, 502, 503, 504 |

Failures inside the gateway return the route's documented envelope, including a failure we did not anticipate. On OpenAI-compatible routes, that arrives as `500` `internal_error` rather than an empty body. The platform edge can answer before the gateway with no JSON envelope: see [Payload too large](#payload-too-large) and [HTML 403 security checkpoint](#html-403-security-checkpoint).

OpenAI-compatible error responses carry `X-Request-Id`, matching `request_id`. Anthropic-compatible responses carry both `request-id` and `x-request-id`. Quote a request id when you contact support.

## Should I retry?

The absence of the retry headers is a deliberate signal, not an oversight. A JSON `400`, `401`, `402`, `403`, `404` or `413` needs a different request or different credentials, so none carries timing. An [HTML 403](#html-403-security-checkpoint) needs a pause before retrying.

Key creation and rotation errors (`api_key_rotation_overlap_limit`, `api_key_mint_busy`) are on [Rotate, revoke, expire](/docs/api-reference/authentication#rotate-revoke-expire).

## Status and code reference

| Status | `error.type` | `error.code` | What you should do |
| - | - | - | - |
| 400 | `invalid_request_error` | `invalid_json` | Send valid JSON. |
| 400 | `invalid_request_error` | `invalid_request` | Fix the field named by `param`. |
| 400 | `invalid_request_error` | `missing_required_field` | Add the field named by `param`. |
| 400 | `invalid_request_error` | `generation_count` | Reduce `n`, `best_of`, or another candidate-count field to four or fewer. |
| 400 | `invalid_request_error` | `context_length_exceeded` | The prompt is longer than the model's context window, and retrying it unchanged fails again. Compact the conversation or send a shorter prompt. `error.context_window` carries the window. On `/v1/messages` the same refusal reads `prompt is too long`; a streaming `/v1/responses` request receives it as a `response.failed` event. Nothing is charged. |
| 400 | `invalid_request_error` | `unsupported_model_operation` | Use an endpoint the selected model supports. |
| 400 | `invalid_request_error` | `hosted_capability_not_supported` | Drop the capability named by `param`, or call a model whose page lists it. |
| 400 | `invalid_request_error` | `hosted_parameter_not_supported` | Fix or remove the field named by `param`. For `response_format`, send the `reasoning_effort` the message names (`none` on most models, `low` on GLM 5.3 Flash), or drop the format. |
| 400 | `invalid_request_error` | `upstream_error` | The model's server refused the request's content, for example a message or parameter shape it cannot serve. Retrying it unchanged fails the same way: fix the field named by `param` when present, otherwise simplify the request. Nothing is charged. |
| 400 | `invalid_request_error` | `invalid_idempotency_key` | Send a non-blank printable ASCII `Idempotency-Key` of 255 characters or fewer. |
| 400 | `invalid_request_error` | `idempotency_mismatch` | This key was already used with a different method, route, or body. Use a new key for a new request. |
| 400 | `invalid_request_error` | `invalid_trace_header` | Fix or remove the header named by `param`. `X-Client-Request-Id` must be non-blank printable ASCII of 512 characters or fewer. |
| 400 | `authentication_error` | `auth_error` | Use a workspace key with `https://api.runinfra.ai/v1`. |
| 401 | `authentication_error` | `invalid_api_key` | Send an active workspace API key in the Bearer header. |
| 402 | `permission_error` | `insufficient_credits` | Add enough credits for the reservation, then retry. |
| 402 | `permission_error` | `account_frozen` | Resolve the billing issue on the workspace, or add credits, then retry. Credits do not lift a hold for a payment dispute: contact support. |
| 402 | `permission_error` | `plan_limit_reached` | The plan is at a limit and neither credits nor Standby can serve the request. Follow `fixes` or wait for the stated reset or refill. |
| 402 | `permission_error` | `spend_limit_reached` | This key reached the monthly spending limit its owner set. Raise or remove the limit under Settings, API keys, or wait for `resets_at`. |
| 403 | `authentication_error` | `auth_error` | Replace a terminal, internal, or deactivated key with an active workspace API key. |
| 403 | `permission_error` | `credits_scope_violation` | The key is scoped to one deployment or endpoint, and `GET /v1/credits` describes the whole workspace. Use a workspace key. |
| 403 | `permission_error` | `usage_scope_violation` | The key is scoped to one deployment or endpoint, and `GET /v1/usage` describes the whole workspace. Use a workspace key. |
| 404 | `not_found_error` | `model_not_found` | Call `GET /v1/models` and use an id returned for the same key. |
| 404 | `not_found_error` | `unknown_endpoint` | The path is not a gateway endpoint, usually a `base_url` that already ends in the endpoint path or a doubled `/v1`. The message names the exact `base_url` to set: OpenAI clients use `https://api.runinfra.ai/v1`, Anthropic clients use the bare host because the SDK appends `/v1/messages` itself. OpenAI-dialect responses also carry `error.base_url` with the value `https://api.runinfra.ai/v1`. |
| 405 | `invalid_request_error` | `method_not_allowed` | The route does not serve this HTTP method. Read the `Allow` header and send one of the methods it lists. |
| 409 | `invalid_request_error` | `idempotency_conflict` | Wait for `Retry-After`, then retry with the same key and body. |
| 413 | `invalid_request_error` | `payload_too_large` | Reduce the JSON request body below 3.5MB. It counts the whole request, including any conversation you resend, and an image is about a third larger once base64 encoded. |
| 422 | `invalid_request_error` | `idempotency_replay_unavailable` | Do not switch to a new idempotency key. Contact support with the request id. |
| 429 | `rate_limit_error` | `rate_limit_exceeded` | Wait for `Retry-After`. The per-key window and the [connection gate](/docs/api-reference/rate-limits#layer-1-the-connection-gate) both send this code. |
| 429 | `rate_limit_error` | `hosted_per_key_concurrency_limit` | Wait for `Retry-After`. It clears when a request using this key finishes. |
| 429 | `rate_limit_error` | `hosted_workspace_concurrency_limit` | Wait for `Retry-After`. It clears when a hosted request in the workspace finishes. |
| 429 | `rate_limit_error` | `hosted_shared_model_concurrency_limit` | Wait for `Retry-After`. Contact support if saturation continues. |
| 429 | `rate_limit_error` | `hosted_plan_concurrency_limit` | Wait for `Retry-After`. It clears when a plan-funded request in the workspace finishes. |
| 429 | `rate_limit_error` | `hosted_standby_workspace_limit` | Wait for `Retry-After`. It clears when a Standby request in the workspace finishes. |
| 429 | `rate_limit_error` | `hosted_standby_aggregate_limit` | Wait for `Retry-After`. It clears when a Standby slot on this model becomes available. |
| 429 | `rate_limit_error` | `hosted_standby_pressure` | Wait for `Retry-After`. It clears when load on the model drops. |
| 429 | `rate_limit_error` | `plan_window_busy` | Your running requests hold the rest of a coding plan limit. The server already waited up to 30 seconds. Retry after `Retry-After`; it clears when one of them finishes. |
| 429 | `rate_limit_error` | `hosted_workspace_tpm_limit` | Wait for `Retry-After`. When one request alone exceeds the limit, waiting cannot help, so reduce its token reservation. |
| 429 | `rate_limit_error` | `hosted_admission_congested` | Wait for `Retry-After`. The gateway shed the request under momentary load. It is not a quota. The wait is randomized between 1 and 4 seconds, up to 8 under heavy load, so refused clients do not return together. |
| 429 | `rate_limit_error` | `hosted_saturated_large_prompt` | The model is momentarily near its concurrency limit and requests with very large prompts are deferred first. Wait for `Retry-After` and retry unchanged; smaller prompts admit normally. It is not a size limit: the same request admits when load drops. |
| 429 | `rate_limit_error` | `hosted_model_at_capacity` | Every replica the request reached (at most two) is at capacity for the moment. Wait for `Retry-After` and retry unchanged; the same request succeeds once a slot frees. It is not a limit on your workspace, it carries no `limit` field and no `X-Hosted-Limit-Scope` header, and nothing is charged. See [the model at capacity](/docs/api-reference/rate-limits#layer-4-the-model-at-capacity). |
| 500 | `api_error` | `internal_error` | Retry after `Retry-After`. Include `X-Request-Id` in a ticket if it persists. |
| 500 | `api_error` | `credit_reservation_failed` | The credit reservation could not be recorded. Retry after `Retry-After`. |
| 502 | `api_error` | `response_format_error` | Retry the request. Contact support if it continues. |
| 502 | `api_error` | `upstream_error` | The model failed while serving the request, after it was admitted. Retry after `Retry-After` when present; otherwise back off briefly. |
| 502 | `api_error` | `upstream_http_error` | The model answered the request with an HTTP failure. Retry after `Retry-After` when present; otherwise back off briefly. |
| 504 | `api_error` | `hosted_inference_response_timeout` | The response exceeded its deadline, and no `Retry-After` is sent. For a non-streaming request, set `stream: true` or lower `max_tokens` before retrying. |
| 200 | `api_error` | `upstream_stream_error` | The stream ended early, on the same HTTP 200: its last frame carries this code, or `upstream_error` or `hosted_inference_response_timeout` when the gateway knows why, and no `[DONE]` follows. Treat a missing `[DONE]` as incomplete and retry. See [Streaming](/docs/api-reference/streaming#frame-shapes). |

## The 503 codes

A `503` means the gateway could not safely serve the request. During a rate-limit store outage, each server applies rate limits on its own instead of returning `limiter_unavailable`, but hosted admission and `Idempotency-Key` checks can still return a retryable `503`. All of them carry `Retry-After` except `upstream_http_error`, which forwards one only when the model supplied it, and `hosted_inference_unavailable`, which carries none. For `datastore_unavailable`, `hosted_evidence_unavailable`, `hosted_admission_unavailable` and the connection gate's `limiter_unavailable`, the wait is randomized between 1 and 4 seconds, up to 8 under heavy load.

| `error.code` | What happened |
| - | - |
| `hosted_model_paused` | The model is temporarily paused. `error.paused_until` carries the next availability check. `Retry-After` is the time until that check, capped at one hour, and 60 seconds once the check time has passed. |
| `limiter_unavailable` | The rate-limit store is not configured, or the connection gate could not evaluate the request, so it was refused rather than served unmetered. |
| `hosted_admission_unavailable` | The capacity admission check could not complete, for example during a rate-limit store outage. |
| `hosted_evidence_unavailable` | The usage record that must exist before serving could not be written. No charge is incurred. |
| `hosted_inference_unavailable` | The model could not be served on this route. Retry shortly; contact support if it persists. |
| `datastore_unavailable` | The record the request depends on could not be read or written. Congestion alone is not this error and answers `429` `hosted_admission_congested` instead. |
| `idempotency_unavailable` | Your `Idempotency-Key` could not be checked, so the request was refused rather than risk running twice. Retry after `Retry-After` with the same key. |
| `billing_settlement_error` | Usage settlement did not complete. |
| `pricing_unavailable` | Pricing for the target could not be resolved. |
| `upstream_http_error` | The model answered the request with an HTTP failure. When the model's side supplied a `Retry-After`, it is forwarded unchanged; if the header is absent, back off briefly before retrying. |
| `upstream_error` | RunInfra could not reach the model for this request and did not complete it. |

## Coding plan limit

A serving coding plan returns `402 plan_limit_reached` when a plan window is at its limit and neither credits nor Standby can serve the request. The response carries `x-should-retry: false` and no `Retry-After`. It is a billing refusal; temporary Standby capacity refusals use `429` with retry timing.

The fields sit directly inside `error`, not under `details`. This illustrative Pro response was constructed with the plan's exact window limits. Credits after limits are disabled and today's Standby is exhausted:

```json wrap theme={"dark"}
{
  "error": {
    "code": "plan_limit_reached",
    "message": "RunInfra coding plan (Pro): 5-hour limit reached, resets Sep 23, 3:00 PM UTC. Standby refills Sep 24, 12:00 AM UTC. Keep full speed on credits: https://runinfra.ai/settings/cost?section=plan&focus=at-limit",
    "limit": "five_hour",
    "resets_at": "2026-09-23T15:00:00.000Z",
    "tier": "pro",
    "used_microcents": 323076922,
    "limit_microcents": 323076922,
    "pending_microcents": 0,
    "standby_resets_at": "2026-09-24T00:00:00.000Z",
    "standby_blocker": "exhausted",
    "credits": {
      "enabled": false,
      "cap_cents": null,
      "balance_cents": 0,
      "month_spent_microcents": 0,
      "pending_max_holds_microcents": 0,
      "blocker": "disabled"
    },
    "fixes": [
      {
        "code": "enable_credits",
        "label": "Keep full speed on credits",
        "url": "https://runinfra.ai/settings/cost?section=plan&focus=at-limit"
      }
    ],
    "type": "permission_error",
    "param": null
  },
  "request_id": "req_example_plan_limit"
}
```

| Field inside `error` | Meaning |
| - | - |
| `limit` | `five_hour` or `week`. If both block, this identifies the blocking window that ends later, when both can admit again. |
| `resets_at` | ISO 8601 reset time for that window. |
| `tier` | `starter`, `pro`, or `team`. |
| `used_microcents`, `limit_microcents`, `pending_microcents` | The blocking window's settled usage, limit, and in-flight commitments. |
| `standby_resets_at` | ISO 8601 daily refill time. A refill does not remove a payment or coverage restriction. |
| `standby_blocker` | Why Standby cannot serve: see below. |
| `credits` | Credits-after-limits choice, cap in cents, balance in cents, spending and pending maximum holds in microcents, and `blocker`. |
| `fixes` | One next step with a machine-readable `code`, display `label`, and `url`, chosen for the blocker and caller's permission. |

| `standby_blocker` | Meaning |
| - | - |
| `not_covered` | Standby covers chat requests only. |
| `payment_needed` | Standby is off until the plan payment goes through. |
| `credits_owed` | Standby is off while the credit balance is negative. |
| `closed` | Standby is unavailable. |
| `exhausted` | The daily amount, including in-flight commitments, has no room. |

`credits.blocker` is `disabled`, `balance_negative`, `key_limit`, `empty`, `cap`, or `period_unavailable`. An owner can receive **Keep full speed on credits**, **Add funds**, **Raise the key limit**, **Raise the cap**, or **Contact support** as the next step. A member receives **Ask an owner**.

The [funding headers](/docs/api-reference/usage#response-funding-headers) report `x-runinfra-funding: refused`. `/v1/messages` rewrites the body to `error.type: "billing_error"` and the message text, with `request_id`; the other plan fields are absent. In Messages clients such as Claude Code, read the message: it carries the limit, reset, Standby condition, and next step without an `error.code`. See [Anthropic Messages](/docs/api-reference/anthropic-messages#coding-plan-funding).

## Envelopes that carry more than the table

<AccordionGroup>
  <Accordion title="402 insufficient_credits, with the balance and the top-up link">
    Credit fields are in cents. These values are an example.

    ```json wrap theme={"dark"}
    {
      "error": {
        "current_balance_cents": 0,
        "required_cents": 1,
        "topup_url": "https://runinfra.ai/settings/cost#credits",
        "message": "Out of credits: balance $0.0000, required $0.0100. Add credits at https://runinfra.ai/settings/cost#credits to continue.",
        "type": "permission_error",
        "param": null,
        "code": "insufficient_credits"
      },
      "request_id": "req_example_0010"
    }
    ```

    Add at least `required_cents - current_balance_cents`, then retry. `topup_url` is absolute.

    When the workspace has a coding plan that is not paying for requests (a payment is overdue, or a payment dispute paused it), `message` ends with the plan's state and its next step, such as paying the overdue amount, or contacting support.

    A `free_window_ended_at` field appears only when the model you called was on a promotional free window that closed within the last seven days, as a UTC RFC 3339 instant describing the same close the `message` states in words. When there is no recently closed window the field is absent rather than null, which is the usual case. The same field accompanies a `402` `account_frozen` under the same rule. A model's price, and whether it is inside a free window, live on that model's page, never here.
  </Accordion>

  <Accordion title="402 spend_limit_reached, with the limit, the spend, and the reset instant">
    A per-key monthly spending limit is optional and set by the key's owner under Settings, API keys. It is enforced on the key's settled spend in the current calendar month (UTC), so a request already in flight when the limit is reached may still complete. Amounts are in cents. These values are an example.

    ```json wrap theme={"dark"}
    {
      "error": {
        "limit_cents": 5000,
        "spent_cents": 5000,
        "resets_at": "2026-10-01T00:00:00.000Z",
        "message": "This API key has reached its monthly spending limit of $50.00 ($50.00 settled this month). Raise or remove the limit under Settings, API keys, or wait for the month to reset at 2026-10-01T00:00:00.000Z.",
        "type": "permission_error",
        "param": null,
        "code": "spend_limit_reached"
      },
      "request_id": "req_example_0011"
    }
    ```

    This describes the pay-as-you-go key limit. On a serving coding plan, the key limit blocks credits after plan limits; eligible plan or Standby usage can still run. See [key payment settings](/docs/api-reference/authentication#coding-plan-and-credits-only). Other keys in the workspace are unaffected. Raise or remove the limit on the key, use another key, or retry after `resets_at`.
  </Accordion>

  <Accordion title="503 hosted_model_paused, with the next availability check">
    ```json wrap theme={"dark"}
    {
      "error": {
        "paused_until": "2026-09-12T11:07:19.211Z",
        "resumeCheckAt": "2026-09-12T11:07:19.211Z",
        "message": "Nemotron 3.5 Lightning 30B is temporarily paused. Availability is rechecked automatically. View https://runinfra.ai/inference-api/nemotron-3-5-lightning-30b for live status, then retry after the interval in the Retry-After header.",
        "type": "api_error",
        "param": null,
        "code": "hosted_model_paused"
      },
      "request_id": "req_example_0014"
    }
    ```

    `Retry-After` is in seconds: the time until the next availability check, capped at one hour. Once that check time has passed, it is 60 seconds on every endpoint, so a client polls about once a minute instead of every second. The status and code stay the same.

    A paused model is not a missing model. It still appears in `GET /v1/models` with `availability: "paused"`, so treat this as retry and poll, never as a reason to drop the model id from your configuration.
  </Accordion>

  <Accordion title="422 idempotency_replay_unavailable, and why not to change the key">
    ```json wrap theme={"dark"}
    {
      "error": {
        "original_status": 200,
        "original_response_bytes": 25000000,
        "message": "The original response for this Idempotency-Key was too large to replay safely. Do not retry with a new key because that can start new paid work. Contact support with the request ID.",
        "type": "invalid_request_error",
        "param": "Idempotency-Key",
        "code": "idempotency_replay_unavailable"
      },
      "request_id": "req_example_0012"
    }
    ```

    ```http theme={"dark"}
    X-RunInfra-Idempotent-Replay: unavailable
    ```

    A new key here would start new inference work and a second charge. Contact support with `request_id` instead.
  </Accordion>

  <Accordion title="400 hosted_parameter_not_supported on response_format">
    ```json wrap theme={"dark"}
    {
      "error": {
        "message": "This model cannot enforce response_format while it is reasoning, so the reply would not match the format you asked for and the request would be billed anyway. Send \"reasoning_effort\": \"none\" together with response_format, or remove response_format. On /v1/responses the same remedy is a top-level \"reasoning_effort\": \"none\" beside \"text\": {\"format\": ...}.",
        "type": "invalid_request_error",
        "param": "response_format",
        "code": "hosted_parameter_not_supported"
      },
      "request_id": "req_example_0020"
    }
    ```

    Send the request again with the effort the message names beside `response_format`, or drop the format. The rule is per model, and the affected models are named in [A JSON format the model cannot hold while reasoning](/docs/api-reference/chat-completions#three-things-a-request-can-be-refused-for). Nothing is charged for the refusal, which is the point: the alternative is a `200` carrying a reply your parser rejects, billed in full.
  </Accordion>

  <Accordion title="429 hosted_admission_congested, a capacity shed">
    Under a momentary burst the gateway waits only briefly for capacity and then sheds the request rather than let it wait longer, so it never completes. That is a temporary capacity condition, not an outage and not a quota, so it answers `429`.

    ```json wrap theme={"dark"}
    {
      "error": {
        "message": "The gateway is at capacity and shed this request before completing it. This is a temporary condition on our side, not a problem with your request. Retry after the interval in the Retry-After header.",
        "type": "rate_limit_error",
        "param": null,
        "code": "hosted_admission_congested"
      },
      "request_id": "req_example_0018"
    }
    ```

    This answer applies on every `/v1` route, so the same condition never surfaces as a different status elsewhere. The message varies with where the shed happened; the status, `type` and `code` do not. Unlike a hosted admission limit, a shed carries no `limit` field and no `X-Hosted-Limit-Scope` header, because there is no quota to size against.
  </Accordion>

  <Accordion title="401 or 403 on a key that looks correct">
    Browser or device sign-in with `runinfra login` creates a terminal key. It has the same `rp_` prefix and 40 character body as a workspace API key and it cannot call the inference API, so it answers `authentication_error` with the message "This API key is internal and cannot be used for customer inference." Create a workspace API key in [Settings, API keys](https://runinfra.ai/settings/api-keys).
  </Accordion>
</AccordionGroup>

## HTML 403 security checkpoint

An HTML `403` security checkpoint instead of a JSON error means automatic traffic protection flagged your network. The response carries `x-vercel-mitigated: challenge`. Wait a few minutes before retrying, lower the request burst, and [contact support](mailto:jaber@runinfra.ai) if it persists.

## Payload too large

The 3.5MB ceiling is enforced on the declared `Content-Length` and again while the body streams, so an understated `Content-Length` does not bypass it. It is a limit on the encoded request and is separate from the model context window: split the prompt across requests or send less context.

Well above the limit, around 4MB and beyond, the platform edge rejects the request before the gateway sees it, with a plain-text `FUNCTION_PAYLOAD_TOO_LARGE` response carrying no JSON envelope and no request id. Treat it as the same instruction: shrink the body.

## Retry rules

* Keep the same `Idempotency-Key` when you retry, streaming or not.
* Respect `Retry-After`, or `Retry-After-Ms` if your client reads it, on every response that carries them.
* On the hosted admission `429` codes the value is an estimate, capped at 8 seconds, and it changes between refusals. Add jitter if many workers share one key.
* Do not replace the key after `422 idempotency_replay_unavailable`.
* A streaming retry never replays the delivered tokens. The same key returns `409` while the original is open, then the terminal usage and cost once it settles. See [Idempotent retries](/docs/api-reference/idempotent-retries).

## Related

<Columns cols={3}>
  <Card title="Rate limits" icon="gauge" href="/docs/api-reference/rate-limits">
    What each 429 means, and how to back off.
  </Card>

  <Card title="Idempotent retries" icon="refresh-cw" href="/docs/api-reference/idempotent-retries">
    Retry safely, streaming or not.
  </Card>

  <Card title="Troubleshooting" icon="life-buoy" href="/docs/tips/troubleshooting">
    Symptom first, from the response you actually got.
  </Card>
</Columns>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.