> ## Documentation Index
> Fetch the complete documentation index at: https://runinfra.ai/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Troubleshooting

> Diagnose Model APIs errors, check your key and balance, and retry safely.

Your coding agent can [set up RunInfra's open models for you](/docs/tools-sdks/agent-setup). For setup issues, see [CLI troubleshooting](/docs/tools-sdks/connect-troubleshooting). For API calls, start with the status code you received.

## An API call was refused

For requests that reach the gateway, OpenAI-compatible routes return `{ error: { message, type, code } }`. `/v1/messages` and `/v1/messages/count_tokens` return the [Anthropic error envelope](/docs/api-reference/anthropic-messages#errors). Both styles include a request id worth quoting in any support ticket. The first question is always whether waiting helps.

For key creation and rotation errors, see [Rotate, revoke, expire](/docs/api-reference/authentication#rotate-revoke-expire).

<AccordionGroup>
  <Accordion title="401 Unauthorized">
    The key is missing, wrong, or expired. Check all three:

    * The header is exactly `Authorization: Bearer YOUR_KEY`, with no extra quotes or whitespace. On `/v1/messages`, `/v1/messages/count_tokens` and the `GET /v1/models` routes, a non-blank `x-api-key` also works and takes precedence.
    * The key is an active workspace API key for Model APIs.
    * The key has not passed its expiry at [Settings > API Keys](https://runinfra.ai/settings/api-keys).

    Create a replacement key in Settings and update your client to use it. On OpenAI-compatible routes, an expired key returns `401 invalid_api_key`.
  </Accordion>

  <Accordion title="402 Payment Required">
    Read the billing reason before changing settings. `insufficient_credits` means the credit balance cannot cover the call: use **Add funds** at [Settings > Billing](https://runinfra.ai/settings/cost#credits). `plan_limit_reached` means the plan is at a limit and neither credits nor Standby can serve this request. Follow the next step in its message or `fixes`, or wait for the stated reset or refill. Plan-limit refusals carry `x-should-retry: false` and no `Retry-After`. See [Plan limit errors](/docs/api-reference/errors#coding-plan-limit). If a settlement drove the balance below zero, your next top-up offsets the negative amount first. See [Pricing and credits](/docs/introduction/plans#at-zero-balance).
  </Accordion>

  <Accordion title="403 Forbidden">
    A deactivated key returns `403` with the message `API key is deactivated`. On OpenAI-compatible routes, `error.code` is `auth_error`. Replace it with an active key.

    Use a workspace API key and select a hosted model with the `model` field. See [Authentication](/docs/api-reference/authentication).

    An HTML security checkpoint needs a pause before retrying. See [HTML 403 security checkpoint](/docs/api-reference/errors#html-403-security-checkpoint).
  </Accordion>

  <Accordion title="404 Not Found">
    The model id does not resolve for your key. Model ids are not OpenAI names, and they are not the title on a model card. Read the live list and pass one of those ids:

    ```bash theme={"dark"}
    curl https://api.runinfra.ai/v1/models \
      -H "Authorization: Bearer $RUNINFRA_GATEWAY_KEY"
    ```
  </Accordion>

  <Accordion title="429 Too Many Requests">
    Read `Retry-After`, wait that long, then send the same request again. Which limit you hit is named in the response headers, and there are two different ones:

    | Headers on the response | Which limit | What changes it |
    | - | - | - |
    | `X-RateLimit-Remaining: 0`, with no `X-Hosted-Limit-Scope` | Your key's requests per minute, over a rolling 60-second window | A key with no custom limit follows the workspace default, so a workspace change applies without rotating the key. A custom per-key limit stays in place, clamped to the workspace maximum. |
    | `X-Hosted-Limit-Scope`, `X-Hosted-Limit-Value` | Hosted admission: concurrency or workspace tokens per minute | Set by the model's capacity and by whether the workspace has bought credit; it applies once the model is busy. Read the current number from `X-Hosted-Limit-Value` or `error.limit` on the response. |

    The window slides, so capacity returns gradually rather than refilling on a clock edge: pacing requests evenly beats bursting. If many workers share one key, add jitter to your retries. A `429` on a key that has never authenticated before is the connection gate, which bounds how fast new credentials can be presented; roll a new key out on one request before fanning out. Read `error.limit` from the response for the current number, and see [Rate limits](/docs/api-reference/rate-limits).

    Your first credit purchase raises the request rate and your share of a busy model. Further purchases raise the token tier. See [Pricing and credits](/docs/introduction/plans#rate-limits).

    Coding plans have separate usage windows and Standby capacity rules. A plan-limit billing refusal is a `402`; a temporary Standby capacity refusal is a `429`. See [Coding plans and Standby](/docs/api-reference/rate-limits#coding-plans-and-standby).
  </Accordion>

  <Accordion title="503 Service Unavailable">
    On a hosted model, `hosted_model_paused` means the model is paused, not gone. The id stays valid and stays listed, `paused_until` carries the next availability check, and nothing is charged. `Retry-After` is the time until that check, capped at one hour, and 60 seconds once the check time has passed. Keep the id in your configuration and retry.

    For other codes, see [The 503 codes](/docs/api-reference/errors#the-503-codes).
  </Accordion>
</AccordionGroup>

## Still stuck

Send feedback from inside the app, or email [support](mailto:jaber@runinfra.ai) with the `X-Request-Id` from the failing response.

## Related

<Columns cols={3}>
  <Card title="Errors" icon="circle-alert" href="/docs/api-reference/errors">
    Every status and code, with the caller action.
  </Card>

  <Card title="Rate limits" icon="gauge" href="/docs/api-reference/rate-limits">
    The four layers that can refuse you.
  </Card>

  <Card title="Idempotent retries" icon="refresh-cw" href="/docs/api-reference/idempotent-retries">
    Retry without a second charge.
  </Card>
</Columns>
