Skip to main content
Your coding agent can set up RunInfra’s open models for you. For setup issues, see CLI troubleshooting. For API calls, start with the status code you received.

An API call was refused

For requests that reach the gateway, OpenAI-compatible routes return { error: { message, type, code } }. /v1/messages and /v1/messages/count_tokens return the Anthropic error envelope. Both styles include a request id worth quoting in any support ticket. The first question is always whether waiting helps. For key creation and rotation errors, see Rotate, revoke, expire.
The key is missing, wrong, or expired. Check all three:
  • The header is exactly Authorization: Bearer YOUR_KEY, with no extra quotes or whitespace. On /v1/messages, /v1/messages/count_tokens and the GET /v1/models routes, a non-blank x-api-key also works and takes precedence.
  • The key is an active workspace API key for Model APIs.
  • The key has not passed its expiry at Settings > API Keys.
Create a replacement key in Settings and update your client to use it. On OpenAI-compatible routes, an expired key returns 401 invalid_api_key.
Read the billing reason before changing settings. insufficient_credits means the credit balance cannot cover the call: use Add funds at Settings > Billing. plan_limit_reached means the plan is at a limit and neither credits nor Standby can serve this request. Follow the next step in its message or fixes, or wait for the stated reset or refill. Plan-limit refusals carry x-should-retry: false and no Retry-After. See Plan limit errors. If a settlement drove the balance below zero, your next top-up offsets the negative amount first. See Pricing and credits.
A deactivated key returns 403 with the message API key is deactivated. On OpenAI-compatible routes, error.code is auth_error. Replace it with an active key.Use a workspace API key and select a hosted model with the model field. See Authentication.An HTML security checkpoint needs a pause before retrying. See HTML 403 security checkpoint.
The model id does not resolve for your key. Model ids are not OpenAI names, and they are not the title on a model card. Read the live list and pass one of those ids:
Read Retry-After, wait that long, then send the same request again. Which limit you hit is named in the response headers, and there are two different ones:The window slides, so capacity returns gradually rather than refilling on a clock edge: pacing requests evenly beats bursting. If many workers share one key, add jitter to your retries. A 429 on a key that has never authenticated before is the connection gate, which bounds how fast new credentials can be presented; roll a new key out on one request before fanning out. Read error.limit from the response for the current number, and see Rate limits.Your first credit purchase raises the request rate and your share of a busy model. Further purchases raise the token tier. See Pricing and credits.Coding plans have separate usage windows and Standby capacity rules. A plan-limit billing refusal is a 402; a temporary Standby capacity refusal is a 429. See Coding plans and Standby.
On a hosted model, hosted_model_paused means the model is paused, not gone. The id stays valid and stays listed, paused_until carries the next availability check, and nothing is charged. Retry-After is the time until that check, capped at one hour, and 60 seconds once the check time has passed. Keep the id in your configuration and retry.For other codes, see The 503 codes.

Still stuck

Send feedback from inside the app, or email support with the X-Request-Id from the failing response.

Errors

Every status and code, with the caller action.

Rate limits

The four layers that can refuse you.

Idempotent retries

Retry without a second charge.