> ## Documentation Index
> Fetch the complete documentation index at: https://runinfra.ai/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Rate limits

> The four layers that can refuse a request with 429, the codes each uses, and how to back off correctly.

Four independent layers, plus [coding plan limits](#coding-plans-and-standby), can refuse a Model APIs request. The four layers run in this order:

1. **A connection gate**, before your credential is read.
2. **A per-key request window**, requests per minute for that key.
3. **Hosted admission**, concurrency and workspace tokens per minute on a shared model.
4. **The model at capacity**, after admission, when every server the request reached is full.

All four answer HTTP `429` with the standard [error envelope](/docs/api-reference/errors). The connection gate and the per-key window both use `rate_limit_exceeded`, hosted admission uses its own codes, and a model at capacity uses `hosted_model_at_capacity`. `/v1/messages` uses the Anthropic envelope, but the same `Retry-After` headers and limits apply.

## Read the header, never a constant

Every `429` carries `Retry-After`. The per-key window, hosted admission and the model at capacity also send `Retry-After-Ms`.

```http theme={"dark"}
Retry-After: <seconds>
Retry-After-Ms: <milliseconds>
```

`Retry-After-Ms` is the same instant in milliseconds. Read it first if your client supports it, and the OpenAI and Anthropic SDKs both do: it is the only one of the two that can express a wait shorter than a second without rounding up. `Retry-After` is the RFC 9110 header in whole seconds, always at least `1`, always rounded up, so it never advises you to retry early.

## Layer 1, the connection gate

Before the API reads your key, a gate bounds how fast unrecognized credentials can be looked up, so a flood of invalid keys cannot crowd out real traffic. It applies per source address and, while the rate-limit store is healthy, across the service as a whole.

There is no warm-up and nothing to request: a new key works on its first call. With a healthy store, a key stops counting against the gate once it succeeds, so steady traffic from a working key never meets it, and rolling a new key out on one request before fanning out avoids it entirely. You are most likely to see it when starting many workers at once with a brand new key, or when a script is retrying a key that is simply wrong.

The budgets are not published, because this is a defensive control and it is tuned. A gate refusal carries no `X-RateLimit-*` headers, which sets it apart from the per-key window.

If the rate-limit store stops answering, each API server applies the same budgets in memory, without sharing counts with other servers. A key the server already trusted keeps bypassing the gate on that server for up to 15 minutes; other keys get that server's new-key budgets.

When the store is not configured or a request cannot be evaluated, the gate returns `503` `limiter_unavailable` rather than admitting unbounded work. Its `Retry-After` is randomized between 1 and 4 seconds, up to 8 seconds under heavy load, so refused clients do not all return at once.

## Layer 2, the per-key request window

A rolling 60 second window while the rate-limit store is healthy. A key with no custom limit follows the workspace's current default, so a workspace upgrade applies without rotating the key. A custom per-key limit stays in place, clamped to the workspace maximum.

```json theme={"dark"}
{
  "error": {
    "message": "Rate limit exceeded. Please retry after the reset time.",
    "type": "rate_limit_error",
    "param": null,
    "code": "rate_limit_exceeded"
  },
  "request_id": "req_example_1000"
}
```

| Header | Meaning |
| - | - |
| `X-RateLimit-Limit` | The key's request limit for the 60 second window. |
| `X-RateLimit-Remaining` | Requests left in the current window. It is `0` on this refusal. |
| `X-RateLimit-Reset` | Unix seconds when the window resets. |
| `Retry-After` | Seconds to wait. |

The window slides, so capacity returns gradually as individual requests age past 60 seconds. There is no clock edge where the whole budget refills at once, which is why pacing evenly beats bursting. During a store outage, each server counts the key's requests in its own fixed 60 second window, and the `X-RateLimit-*` headers describe that server's count. Going over the limit there returns `429`. If the store is not configured, the request returns `503` `limiter_unavailable` rather than being served unmetered.

Hosted admission and requests with an `Idempotency-Key` still need the store, so they can return a retryable `503` during an outage. See [the 503 codes](/docs/api-reference/errors#the-503-codes).

## Layer 3, hosted admission

<div className="block dark:hidden">
  <svg viewBox="0 0 720 374" width="100%" role="img" aria-label="The hosted admission gates in check order, where each limit comes from, and the 429 headers that name the gate you hit." fill="none" xmlns="http://www.w3.org/2000/svg"><text x="24" y="16" fill="#6e6d64" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">Hosted admission</text><text x="696" y="16" fill="#78786f" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0" textAnchor="end">checked in this order</text><rect x="24.5" y="32.5" width="209" height="63" fill="#ffffff" stroke="#e8e8e3" strokeWidth="1" shapeRendering="crispEdges" /><text x="38" y="50" fill="#6e6d64" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">Model capacity</text><text x="38" y="70" fill="#0f0f0e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="12.5" fontWeight="500" letterSpacing="-0.12">Configured per model</text><text x="38" y="85" fill="#6e6d64" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">published in GET /v1/models</text><line x1="234" y1="64" x2="272" y2="64" stroke="#e8e8e3" strokeWidth="1" strokeDasharray="3 3" /><line x1="272" y1="30" x2="272" y2="304" stroke="#e8e8e3" strokeWidth="1" strokeDasharray="3 3" /><rect x="269.5" y="50.5" width="5" height="5" fill="#bbb9b1" shapeRendering="crispEdges" /><rect x="286.5" y="30.5" width="409" height="45" fill="#ffffff" stroke="#e8e8e3" strokeWidth="1" shapeRendering="crispEdges" /><text x="300" y="48" fill="#6e6d64" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">1  API key</text><text x="300" y="65" fill="#6e6d64" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">the busy-model share; more while idle</text><rect x="269.5" y="106.5" width="5" height="5" fill="#bbb9b1" shapeRendering="crispEdges" /><rect x="286.5" y="86.5" width="409" height="45" fill="#ffffff" stroke="#e8e8e3" strokeWidth="1" shapeRendering="crispEdges" /><text x="300" y="104" fill="#6e6d64" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">2  Workspace</text><text x="300" y="121" fill="#6e6d64" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">the busy-model share; more while idle</text><rect x="269.5" y="162.5" width="5" height="5" fill="#bbb9b1" shapeRendering="crispEdges" /><rect x="286.5" y="142.5" width="409" height="45" fill="#ffffff" stroke="#e8e8e3" strokeWidth="1" shapeRendering="crispEdges" /><text x="300" y="160" fill="#6e6d64" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">3  Large-prompt shed</text><text x="682" y="160" fill="#5a8f00" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0" textAnchor="end">not a size limit</text><text x="300" y="177" fill="#6e6d64" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">very large prompts only, near full capacity</text><rect x="269.5" y="218.5" width="5" height="5" fill="#bbb9b1" shapeRendering="crispEdges" /><rect x="286.5" y="198.5" width="409" height="45" fill="#ffffff" stroke="#e8e8e3" strokeWidth="1" shapeRendering="crispEdges" /><text x="300" y="216" fill="#6e6d64" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">4  Shared model</text><text x="300" y="233" fill="#6e6d64" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">the full capacity, all workspaces</text><rect x="269.5" y="274.5" width="5" height="5" fill="#bbb9b1" shapeRendering="crispEdges" /><rect x="286.5" y="254.5" width="409" height="45" fill="#ffffff" stroke="#e8e8e3" strokeWidth="1" shapeRendering="crispEdges" /><text x="300" y="272" fill="#6e6d64" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">5  Workspace tokens per minute</text><text x="300" y="289" fill="#6e6d64" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">per-model budget, rolling 60s window</text><line x1="24" y1="320" x2="696" y2="320" stroke="#e8e8e3" strokeWidth="1" strokeDasharray="3 3" /><rect x="21.5" y="317.5" width="5" height="5" fill="#bbb9b1" shapeRendering="crispEdges" /><rect x="693.5" y="317.5" width="5" height="5" fill="#bbb9b1" shapeRendering="crispEdges" /><rect x="24" y="334" width="5" height="5" fill="#b07f24" shapeRendering="crispEdges" /><text x="36" y="342" fill="#b07f24" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0">429</text><text x="86" y="342" fill="#6e6d64" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">with Retry-After, X-Hosted-Limit-Scope, and X-Hosted-Limit-Value naming the gate</text><text x="36" y="360" fill="#78786f" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">Read error.limit and X-Hosted-Limit-Value from the response; values change with capacity.</text></svg>
</div>

<div className="hidden dark:block">
  <svg viewBox="0 0 720 374" width="100%" role="img" aria-label="The hosted admission gates in check order, where each limit comes from, and the 429 headers that name the gate you hit." fill="none" xmlns="http://www.w3.org/2000/svg"><text x="24" y="16" fill="#8f8e83" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">Hosted admission</text><text x="696" y="16" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0" textAnchor="end">checked in this order</text><rect x="24.5" y="32.5" width="209" height="63" fill="#161614" stroke="#383833" strokeWidth="1" shapeRendering="crispEdges" /><text x="38" y="50" fill="#8f8e83" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">Model capacity</text><text x="38" y="70" fill="#f0efe2" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="12.5" fontWeight="500" letterSpacing="-0.12">Configured per model</text><text x="38" y="85" fill="#9a998e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">published in GET /v1/models</text><line x1="234" y1="64" x2="272" y2="64" stroke="#383833" strokeWidth="1" strokeDasharray="3 3" /><line x1="272" y1="30" x2="272" y2="304" stroke="#383833" strokeWidth="1" strokeDasharray="3 3" /><rect x="269.5" y="50.5" width="5" height="5" fill="#6e6d64" shapeRendering="crispEdges" /><rect x="286.5" y="30.5" width="409" height="45" fill="#161614" stroke="#383833" strokeWidth="1" shapeRendering="crispEdges" /><text x="300" y="48" fill="#8f8e83" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">1  API key</text><text x="300" y="65" fill="#9a998e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">the busy-model share; more while idle</text><rect x="269.5" y="106.5" width="5" height="5" fill="#6e6d64" shapeRendering="crispEdges" /><rect x="286.5" y="86.5" width="409" height="45" fill="#161614" stroke="#383833" strokeWidth="1" shapeRendering="crispEdges" /><text x="300" y="104" fill="#8f8e83" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">2  Workspace</text><text x="300" y="121" fill="#9a998e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">the busy-model share; more while idle</text><rect x="269.5" y="162.5" width="5" height="5" fill="#6e6d64" shapeRendering="crispEdges" /><rect x="286.5" y="142.5" width="409" height="45" fill="#161614" stroke="#383833" strokeWidth="1" shapeRendering="crispEdges" /><text x="300" y="160" fill="#8f8e83" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">3  Large-prompt shed</text><text x="682" y="160" fill="#8fd400" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0" textAnchor="end">not a size limit</text><text x="300" y="177" fill="#9a998e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">very large prompts only, near full capacity</text><rect x="269.5" y="218.5" width="5" height="5" fill="#6e6d64" shapeRendering="crispEdges" /><rect x="286.5" y="198.5" width="409" height="45" fill="#161614" stroke="#383833" strokeWidth="1" shapeRendering="crispEdges" /><text x="300" y="216" fill="#8f8e83" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">4  Shared model</text><text x="300" y="233" fill="#9a998e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">the full capacity, all workspaces</text><rect x="269.5" y="274.5" width="5" height="5" fill="#6e6d64" shapeRendering="crispEdges" /><rect x="286.5" y="254.5" width="409" height="45" fill="#161614" stroke="#383833" strokeWidth="1" shapeRendering="crispEdges" /><text x="300" y="272" fill="#8f8e83" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">5  Workspace tokens per minute</text><text x="300" y="289" fill="#9a998e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">per-model budget, rolling 60s window</text><line x1="24" y1="320" x2="696" y2="320" stroke="#383833" strokeWidth="1" strokeDasharray="3 3" /><rect x="21.5" y="317.5" width="5" height="5" fill="#6e6d64" shapeRendering="crispEdges" /><rect x="693.5" y="317.5" width="5" height="5" fill="#6e6d64" shapeRendering="crispEdges" /><rect x="24" y="334" width="5" height="5" fill="#d9a64a" shapeRendering="crispEdges" /><text x="36" y="342" fill="#d9a64a" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0">429</text><text x="86" y="342" fill="#9a998e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">with Retry-After, X-Hosted-Limit-Scope, and X-Hosted-Limit-Value naming the gate</text><text x="36" y="360" fill="#9a998e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">Read error.limit and X-Hosted-Limit-Value from the response; values change with capacity.</text></svg>
</div>

The gates run in the order of the table below. Each one counts per model. A hosted refusal carries `limit` inside `error`, plus headers naming the gate that stopped you.

```http theme={"dark"}
X-Hosted-Limit-Scope: <scope>
X-Hosted-Limit-Value: <limit>
Retry-After: <seconds>
Retry-After-Ms: <milliseconds>
```

| `error.code` | `X-Hosted-Limit-Scope` | What clears it |
| - | - | - |
| `hosted_per_key_concurrency_limit` | `api-key-concurrency` | An in-flight request using the same key finishes, or load on the model drops. The allowance in `error.limit` applies once the model is busy; with capacity to spare, a key is admitted past it, up to a higher ceiling, and a refusal there quotes that ceiling instead. |
| `hosted_workspace_concurrency_limit` | `workspace-concurrency` | Any in-flight hosted request in the workspace finishes, or load on the model drops. Same rule as the key gate. |
| `hosted_saturated_large_prompt` | `shared-model-saturation` | Load on the model drops, usually within seconds. Applies only to very large prompts while the model is near its concurrency limit; it is not a size limit and the same request admits unchanged when load drops. |
| `hosted_shared_model_concurrency_limit` | `shared-model-concurrency` | Shared capacity on that model frees up. |
| `hosted_workspace_tpm_limit` | `workspace-tokens-per-minute` | The token window rolls forward, or you send fewer tokens. |
| `hosted_admission_congested` | none | A momentary burst; wait for `Retry-After`. Not a quota. |

Every hosted refusal uses the same envelope. This is the API-key concurrency case, and the numbers in it are an example:

```json wrap theme={"dark"}
{
  "error": {
    "limit": 16,
    "message": "This workspace reached its limit of 16 concurrent hosted requests for this model. API keys share this allowance. Retry after the interval in the Retry-After header, or send fewer concurrent requests across the workspace.",
    "type": "rate_limit_error",
    "param": null,
    "code": "hosted_per_key_concurrency_limit"
  },
  "request_id": "req_example_1001"
}
```

Always read `error.limit` and `X-Hosted-Limit-Value` from the response you got, never a number from this page. They change with capacity.

Hosted admission `429` responses also carry `X-RateLimit-Limit-Tokens`, `X-RateLimit-Remaining-Tokens`, and `X-RateLimit-Reset-Tokens` when the limiter has a measured workspace token snapshot. This includes concurrency refusals, so a client can back off with both the refusing concurrency limit and the remaining token budget in view.

## Layer 4, the model at capacity

A request can pass every gate above and still find the model's servers full. When the server a request reaches is at capacity, the gateway tries one other server if the model has more than one. If every server it reached was full, the request answers `429` `rate_limit_error` with code `hosted_model_at_capacity` and a `Retry-After` header.

Wait for `Retry-After` and retry unchanged; the same request succeeds once a slot frees. It is not a limit on your workspace, so it carries no `limit` field and no `X-Hosted-Limit-Scope` header, and nothing is charged. An SDK with retries enabled handles it like any other `429`.

## Handling a 429 in code

<CodeGroup>
  ```python Python theme={"dark"}
  import os
  from openai import OpenAI

  client = OpenAI(
      base_url="https://api.runinfra.ai/v1",
      api_key=os.environ["RUNINFRA_GATEWAY_KEY"],
      max_retries=5,
  )

  response = client.chat.completions.create(
      model="nemotron-3-5-lightning-30b",
      messages=[{"role": "user", "content": "Health check"}],
  )
  print(response.choices[0].message.content)
  ```

  ```typescript TypeScript theme={"dark"}
  import OpenAI from "openai";

  const client = new OpenAI({
    baseURL: "https://api.runinfra.ai/v1",
    apiKey: process.env.RUNINFRA_GATEWAY_KEY,
    maxRetries: 5,
  });

  const response = await client.chat.completions.create({
    model: "nemotron-3-5-lightning-30b",
    messages: [{ role: "user", content: "Health check" }],
  });
  console.log(response.choices[0]?.message?.content);
  ```
</CodeGroup>

An SDK with retries enabled does the right thing with no code from you, because every layer sends `Retry-After`. If you back off yourself, add jitter when many workers share a key so they do not all return at the same instant.

## Three cases worth planning for

* **A `429` before any request of yours has succeeded**, on a key you just created, is the connection gate rather than a quota. Retry once after `Retry-After` with a single request, then resume normal concurrency.
* **One request that alone exceeds the TPM limit** is the case where waiting cannot help. Its refusal still carries `Retry-After`, but retrying it unchanged fails again. Reduce the input or the maximum output tokens before retrying. Hosted token-rate admission uses the same request token estimate as billing, not a byte count.
* **Adding API keys does not buy capacity.** Every key in a workspace draws on one shared concurrency allowance, so a second key does not raise it. Read the current allowance for a model from `max_concurrent_requests_per_api_key` in `GET /v1/models` rather than assuming a fixed figure.

## Coding plans and Standby

A coding plan's 5-hour and weekly limits meter usage separately from request and capacity limits. Exhausting plan funding returns a [402 plan limit](/docs/api-reference/errors#coding-plan-limit), with `x-should-retry: false` and no `Retry-After`.

Requests still running reserve their share of a limit until they finish. When those reservations alone hold the rest of a 5-hour or weekly limit, the credits cap, or today's Standby, a new request waits on the server for up to 30 seconds. If it is still held, the request gets `429 plan_window_busy` with `Retry-After` and `x-should-retry: true`. `error.details.window` is `5h`, `weekly`, `credits_cap`, or `standby_day`. The request is not charged and never moves to credits or Standby.

While a plan serves covered requests, Pro and Team have a temporary floor of at least 20 million tokens per minute per model. Net plan payments contribute to the earned token tier. The temporary floor ends when the plan stops serving.

Plan-funded requests have their own concurrency allowance per model, published as `max_concurrent_plan_requests_per_workspace` in [`GET /v1/models`](/docs/api-reference/models). Standby serves chat only and depends on model capacity. A temporary plan or Standby refusal uses these `429` codes and includes retry timing:

| `error.code` | `X-Hosted-Limit-Scope` | What clears it |
| - | - | - |
| `hosted_plan_concurrency_limit` | `plan-workspace-concurrency` | A plan-funded request in the workspace finishes. |
| `hosted_standby_workspace_limit` | `standby-workspace-concurrency` | A Standby request in the workspace finishes. |
| `hosted_standby_aggregate_limit` | `standby-model-concurrency` | A Standby slot on this model becomes available. |
| `hosted_standby_pressure` | `standby-model-pressure` | Load on the model drops. |

Read the applied value from `error.limit` or `X-Hosted-Limit-Value`. A daily Standby amount does not guarantee that the model can admit a request now.

## Related

<Columns cols={2}>
  <Card title="Errors" icon="circle-alert" href="/docs/api-reference/errors">
    Every status, code, and whether a retry can help.
  </Card>

  <Card title="Models" icon="list" href="/docs/api-reference/models">
    Discover the limits that bind your key at boot.
  </Card>
</Columns>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.