Skip to main content
Setup issues: see CLI troubleshooting, or let your coding agent set up RunInfra. Key creation and rotation errors: see Rotate, revoke, expire.

An API call was refused

For requests that reach the gateway, OpenAI-compatible routes return { error: { message, type, code } }. /v1/messages and /v1/messages/count_tokens return the Anthropic error envelope. Both styles include a request id.
The key is missing, wrong, or expired. Check all three:
  • The header is exactly Authorization: Bearer YOUR_KEY, with no extra quotes or whitespace. On /v1/messages, /v1/messages/count_tokens and the GET /v1/models routes, a non-blank x-api-key also works and takes precedence.
  • The key is an active workspace API key for Model APIs.
  • The key has not passed its expiry at Settings > API keys.
Create a replacement key in Settings and update your client to use it. On OpenAI-compatible routes, an expired key returns 401 invalid_api_key.
Read the billing reason before changing settings:
  • insufficient_credits: the balance cannot cover the call. Use Add funds at Settings > Billing. If a settlement drove the balance below zero, your next top-up offsets it first. See Pricing and credits.
  • spend_limit_reached: the key reached the monthly spending limit its owner set. Raise or remove it at Settings > API keys, or wait for resets_at.
  • account_frozen: resolve the billing issue, or add credits. Credits do not lift a payment dispute hold: contact support.
  • plan_limit_reached: the plan is at a limit and neither credits nor Standby can serve this request. Follow the next step in its message or fixes, or wait for the stated reset or refill. It carries x-should-retry: false and no Retry-After. See Coding plan limit.
A deactivated key returns 403 with the message API key is deactivated. On OpenAI-compatible routes, error.code is auth_error. Replace it with an active key.A terminal key from runinfra login has the same rp_ prefix as a workspace API key but cannot call the API. Use a workspace API key from Settings > API keys and select the model with the model field. See Authentication.An HTML security checkpoint needs a pause before retrying. See HTML 403 security checkpoint.
model_not_found: the model id does not resolve for your key. Model ids are not OpenAI names, and they are not the title on a model card. Read the live list and pass one of those ids:
unknown_endpoint: the path is not a gateway endpoint, usually a doubled /v1 in base_url. The message names the base_url to set.
Read Retry-After, wait that long, then send the same request again. The headers name the limit you hit:
  • X-RateLimit-Remaining: 0, no X-Hosted-Limit-Scope: your key’s requests per minute, over a rolling 60-second window. A key with no custom limit follows the workspace default, so a workspace change applies without rotating the key. A custom per-key limit stays in place, clamped to the workspace maximum.
  • X-Hosted-Limit-Scope and X-Hosted-Limit-Value: hosted admission, concurrency or workspace tokens per minute. It is set by the model’s capacity and by whether the workspace has bought credit, and it applies once the model is busy. The current number is in X-Hosted-Limit-Value or error.limit.
The window slides, so capacity returns gradually rather than refilling on a clock edge: pacing requests evenly beats bursting. If many workers share one key, add jitter to your retries. A 429 on a key that has never authenticated before is the connection gate, which bounds how fast new credentials can be presented; roll a new key out on one request before fanning out. See Rate limits.Your first credit purchase raises the request rate and your share of a busy model. Further purchases raise the token tier. See Pricing and credits.Coding plans have separate usage windows and Standby capacity rules. A plan-limit billing refusal is a 402; a temporary Standby capacity refusal is a 429. See Coding plans and Standby.
On a hosted model, hosted_model_paused means the model is paused, not gone. The id stays valid and stays listed, paused_until carries the next availability check, and nothing is charged. Retry-After is the time until that check, capped at one hour, and 60 seconds once the check time has passed. Keep the id in your configuration and retry.For other codes, see The 503 codes.

Still stuck

Use Send feedback, the RunInfra logo button at the bottom right of the dashboard, or email support with the X-Request-Id from the failing response.

Errors

Every status and code, with the caller action.

Rate limits

The four layers that can refuse you.

Idempotent retries

Retry without a second charge.