> ## Documentation Index
> Fetch the complete documentation index at: https://runinfra.ai/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# List models

> Discover the model ids a workspace key can reach, and the limits the gateway enforces on them.

```http theme={"dark"}
GET https://api.runinfra.ai/v1/models
```

This is how an application discovers, at boot, which model ids its key can call and what binds them. Never hard-code a list you guessed.

## Request

<CodeGroup>
  ```python Python theme={"dark"}
  import os
  from openai import OpenAI

  client = OpenAI(
      base_url="https://api.runinfra.ai/v1",
      api_key=os.environ["RUNINFRA_GATEWAY_KEY"],
  )
  for model in client.models.list().data:
      print(model.id)
  ```

  ```typescript TypeScript theme={"dark"}
  import OpenAI from "openai";

  const client = new OpenAI({
    baseURL: "https://api.runinfra.ai/v1",
    apiKey: process.env.RUNINFRA_GATEWAY_KEY,
  });
  const models = await client.models.list();
  console.log(models.data.map((model) => model.id));
  ```

  ```bash cURL theme={"dark"}
  curl https://api.runinfra.ai/v1/models \
    -H "Authorization: Bearer $RUNINFRA_GATEWAY_KEY"
  ```
</CodeGroup>

## Response

```json theme={"dark"}
{
  "object": "list",
  "data": [
    {
      "id": "nemotron-3-5-lightning-30b",
      "object": "model",
      "owned_by": "runinfra",
      "created": 1786619239,
      "availability": "available",
      "max_request_bytes": 3670016,
      "context_window": 262144,
      "context_length": 262144,
      "max_output_tokens": 262144,
      "pricing": { "input": 0.05, "output": 0.15 },
      "cached_input_price_usd_per_mtok": 0.01,
      "cached_input_price": 0.01,
      "cache_isolation": "isolated",
      "cache_tiers": ["gpu"],
      "cache_retention": "best_effort",
      "max_concurrent_requests_per_api_key": 16,
      "max_tokens_per_minute_per_workspace": 4000000
    }
  ]
}
```

An abbreviated example. Read the values from your own response; they move with capacity.

| Field | Type | Value |
| - | - | - |
| `object` | string | `list` at the top level, `model` for each item. |
| `data[].id` | string | The id to send in the `model` field of an inference request. |
| `data[].owned_by` | string | `runinfra`. |
| `data[].created` | integer | Unix seconds, the creation time of the underlying record. |
| `data[].display_name` | string | The model's display name, such as `Nemotron 3.5 Lightning 30B`. |
| `data[].supported_endpoints` | string\[] | Endpoints that serve this id: `chat_completions`, `responses` and `messages` for a chat model. |
| `data[].modality` | string | `llm` for a chat model. |
| `data[].tool_calling` | boolean | `true` when the model accepts `tools`. |
| `data[].json_mode` | boolean | `true` when the model accepts a JSON `response_format`. |
| `data[].input_modalities` | string\[] | `["text"]`, or `["text", "image"]` when the model accepts images. |
| `data[].reasoning_efforts` | string\[] | `reasoning_effort` values measured to behave distinctly on this model. Empty when none are declared. |
| `data[].default_reasoning_effort` | string or null | The measured default when you omit `reasoning_effort`. `null` means none is declared, not that reasoning is off. |
| `data[].max_completion_tokens` | integer or null | The same value as `max_output_tokens`, under the OpenAI request field's name. `null` when no output limit is declared. |
| `data[].availability` | string | `available` or `paused`. |
| `data[].max_request_bytes` | integer | Largest request body the gateway accepts. Enforced at the gateway, so it is the same for every model. |
| `data[].context_window` | integer | Tokens the model can hold in one request. Present only where a served window has been reported. |
| `data[].context_length` | integer | The same served window under the name OpenAI-compatible catalog readers commonly use. |
| `data[].max_output_tokens` | integer | Generated tokens allowed in one request. Oversized output budgets are clamped to the model's remaining context and reported with `X-RunInfra-Output-Clamped: true` and `X-RunInfra-Output-Token-Ceiling`. |
| `data[].response_time_ceiling_seconds` | integer | Wall-clock seconds one response may run. See [Output limits](/docs/api-reference/chat-completions#output-limits). |
| `data[].pricing` | object | Hosted input and output prices in USD per one million tokens. |
| `data[].cached_input_price_usd_per_mtok` | number | Hosted cached-input price in USD per one million tokens, when the effective cached tier is active. |
| `data[].cached_input_price` | number | The same cached-input price in USD per one million tokens under the short discovery key. |
| `data[].cache_isolation` | string | `isolated` when the active request path separates cache entries by tenant, otherwise `shared`. |
| `data[].cache_tiers` | string\[] | Proven storage tiers, such as `gpu` or `host_ram`. |
| `data[].cache_retention` | string | `best_effort` when cached prefixes are kept as long as capacity allows and may be evicted under memory pressure, or `tiered_host_memory` when the proven cache also spans host memory. Both are best effort and no TTL is implied. The human-readable statement is on the model's page. |
| `data[].max_concurrent_requests_per_api_key` | integer | Requests one key may have in flight against this model once the model is busy. A workspace resolves to the same number. While the model has capacity to spare, requests beyond it are admitted; see [rate limits](/docs/api-reference/rate-limits). |
| `data[].max_concurrent_plan_requests_per_workspace` | integer | Present only while the workspace's coding plan serves a model it covers. Requests the plan funds, on the plan or on credits after its limits, may run this many at once across the workspace once the model is busy. While the model has capacity to spare the gateway may admit more, so this understates and never overstates the allowance. |
| `data[].max_concurrent_standby_requests_per_workspace` | integer | Present beside the field above on a chat model. Standby requests the workspace may run at once, one allowance shared across every model it calls on Standby. `0` means Standby is closed on this model. |
| `data[].max_tokens_per_minute_per_workspace` | integer | Tokens per minute one workspace may spend on this model, over a rolling 60 second window. |
| `data[].session_affinity_idle_ttl_seconds` | integer | Idle seconds before a session's replica placement is forgotten. See [Session affinity](/docs/api-reference/session-affinity). |
| `data[].paused_until` | string | ISO 8601 time availability is next rechecked, not a promised return. Present only when `availability` is `paused`. |

## Notes

* The limit fields exist so you never learn a limit from a `429`. They are read from the code that enforces them, so a published limit and an applied limit cannot drift apart.
* These fields are additive to and disjoint from the fields the OpenAI Model object declares, so a client typed against the OpenAI SDK keeps parsing this response and ignores what it does not recognize.
* Absent means undeclared, never zero and never unlimited. A field appears only where the limit it names is actually enforced for that model; a model that has not reported a served window carries no `context_window` rather than borrowing another model's number.
* `max_request_bytes` is a byte ceiling, not a token ceiling, and the two bind independently. For a very long prompt the byte ceiling is often the one that binds first, whatever the context window says.
* Hosted model prices are returned here for discovery and repeated on each model's page in the [Model Library](https://runinfra.ai/inference-api), with the human-readable cache eviction and retention statement, served precision when verified, capabilities, and the recommended output budget.
* A workspace key lists every public hosted model, available or paused.

## Available and paused

<div className="block dark:hidden">
  <svg viewBox="0 0 720 214" width="100%" role="img" aria-label="The two availability states a buyer sees: available serves 200 responses, paused answers 503 hosted_model_paused with its next availability check, and the model stays listed in both." fill="none" xmlns="http://www.w3.org/2000/svg"><text x="24" y="16" fill="#6e6d64" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">Model availability</text><text x="696" y="16" fill="#78786f" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0" textAnchor="end">GET /v1/models</text><rect x="24.5" y="34.5" width="311" height="115" fill="#ffffff" stroke="#e8e8e3" strokeWidth="1" shapeRendering="crispEdges" /><rect x="38" y="45" width="94" height="3" fill="#76b900" shapeRendering="crispEdges" /><rect x="38" y="48" width="104" height="22" fill="#76b900" shapeRendering="crispEdges" /><rect x="48" y="70" width="94" height="3" fill="#76b900" shapeRendering="crispEdges" /><text x="90" y="63" fill="#0f0f0e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="12" fontWeight="500" textAnchor="middle" letterSpacing="-0.12">available</text><text x="38" y="94" fill="#52524c" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0">availability: available</text><rect x="38" y="108" width="5" height="5" fill="#76b900" shapeRendering="crispEdges" /><text x="50" y="116" fill="#5a8f00" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0">200</text><text x="84" y="116" fill="#6e6d64" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">requests are served and billed</text><text x="38" y="136" fill="#78786f" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">Listed with its id, prices on its page.</text><rect x="384.5" y="34.5" width="311" height="115" fill="#ffffff" stroke="#e8e8e3" strokeWidth="1" shapeRendering="crispEdges" /><rect x="398" y="45" width="78" height="3" fill="#b07f24" shapeRendering="crispEdges" /><rect x="398" y="48" width="88" height="22" fill="#b07f24" shapeRendering="crispEdges" /><rect x="408" y="70" width="78" height="3" fill="#b07f24" shapeRendering="crispEdges" /><text x="442" y="63" fill="#0f0f0e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="12" fontWeight="500" textAnchor="middle" letterSpacing="-0.12">paused</text><text x="398" y="94" fill="#52524c" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0">availability: paused</text><rect x="398" y="108" width="5" height="5" fill="#b07f24" shapeRendering="crispEdges" /><text x="410" y="116" fill="#b07f24" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0">503</text><text x="444" y="116" fill="#78786f" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0">hosted\_model\_paused</text><text x="398" y="136" fill="#78786f" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">Nothing is charged. Honor retry headers.</text><line x1="336" y1="82" x2="384" y2="82" stroke="#bbb9b1" strokeWidth="1" /><path d="M 384 82 l -5 -3 v 6 z" fill="#bbb9b1" /><line x1="384" y1="104" x2="336" y2="104" stroke="#bbb9b1" strokeWidth="1" /><path d="M 336 104 l 5 -3 v 6 z" fill="#bbb9b1" /><text x="360" y="76" fill="#78786f" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0" textAnchor="middle">pause</text><text x="360" y="118" fill="#78786f" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0" textAnchor="middle">return</text><line x1="24" y1="168" x2="696" y2="168" stroke="#e8e8e3" strokeWidth="1" strokeDasharray="3 3" /><rect x="21.5" y="165.5" width="5" height="5" fill="#bbb9b1" shapeRendering="crispEdges" /><rect x="693.5" y="165.5" width="5" height="5" fill="#bbb9b1" shapeRendering="crispEdges" /><text x="24" y="186" fill="#6e6d64" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">A paused model is not a missing model: the id stays valid and stays listed, with paused\_until</text><text x="24" y="200" fill="#6e6d64" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">carrying the next availability check. Poll and retry; never drop the id from your configuration.</text></svg>
</div>

<div className="hidden dark:block">
  <svg viewBox="0 0 720 214" width="100%" role="img" aria-label="The two availability states a buyer sees: available serves 200 responses, paused answers 503 hosted_model_paused with its next availability check, and the model stays listed in both." fill="none" xmlns="http://www.w3.org/2000/svg"><text x="24" y="16" fill="#8f8e83" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">Model availability</text><text x="696" y="16" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0" textAnchor="end">GET /v1/models</text><rect x="24.5" y="34.5" width="311" height="115" fill="#161614" stroke="#383833" strokeWidth="1" shapeRendering="crispEdges" /><rect x="38" y="45" width="94" height="3" fill="#76b900" shapeRendering="crispEdges" /><rect x="38" y="48" width="104" height="22" fill="#76b900" shapeRendering="crispEdges" /><rect x="48" y="70" width="94" height="3" fill="#76b900" shapeRendering="crispEdges" /><text x="90" y="63" fill="#0f0f0e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="12" fontWeight="500" textAnchor="middle" letterSpacing="-0.12">available</text><text x="38" y="94" fill="#c8c7ba" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0">availability: available</text><rect x="38" y="108" width="5" height="5" fill="#76b900" shapeRendering="crispEdges" /><text x="50" y="116" fill="#8fd400" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0">200</text><text x="84" y="116" fill="#9a998e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">requests are served and billed</text><text x="38" y="136" fill="#9a998e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">Listed with its id, prices on its page.</text><rect x="384.5" y="34.5" width="311" height="115" fill="#161614" stroke="#383833" strokeWidth="1" shapeRendering="crispEdges" /><rect x="398" y="45" width="78" height="3" fill="#d9a64a" shapeRendering="crispEdges" /><rect x="398" y="48" width="88" height="22" fill="#d9a64a" shapeRendering="crispEdges" /><rect x="408" y="70" width="78" height="3" fill="#d9a64a" shapeRendering="crispEdges" /><text x="442" y="63" fill="#0f0f0e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="12" fontWeight="500" textAnchor="middle" letterSpacing="-0.12">paused</text><text x="398" y="94" fill="#c8c7ba" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0">availability: paused</text><rect x="398" y="108" width="5" height="5" fill="#d9a64a" shapeRendering="crispEdges" /><text x="410" y="116" fill="#d9a64a" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0">503</text><text x="444" y="116" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0">hosted\_model\_paused</text><text x="398" y="136" fill="#9a998e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">Nothing is charged. Honor retry headers.</text><line x1="336" y1="82" x2="384" y2="82" stroke="#6e6d64" strokeWidth="1" /><path d="M 384 82 l -5 -3 v 6 z" fill="#6e6d64" /><line x1="384" y1="104" x2="336" y2="104" stroke="#6e6d64" strokeWidth="1" /><path d="M 336 104 l 5 -3 v 6 z" fill="#6e6d64" /><text x="360" y="76" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0" textAnchor="middle">pause</text><text x="360" y="118" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0" textAnchor="middle">return</text><line x1="24" y1="168" x2="696" y2="168" stroke="#383833" strokeWidth="1" strokeDasharray="3 3" /><rect x="21.5" y="165.5" width="5" height="5" fill="#6e6d64" shapeRendering="crispEdges" /><rect x="693.5" y="165.5" width="5" height="5" fill="#6e6d64" shapeRendering="crispEdges" /><text x="24" y="186" fill="#9a998e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">A paused model is not a missing model: the id stays valid and stays listed, with paused\_until</text><text x="24" y="200" fill="#9a998e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">carrying the next availability check. Poll and retry; never drop the id from your configuration.</text></svg>
</div>

A paused model keeps its place in the list. It is a real configured model that is temporarily not serving, so poll this endpoint or retry the call itself to detect it coming back. The id does not change across a pause. While paused, the row omits `pricing`, the cached-input price fields, `max_concurrent_requests_per_api_key` and `max_tokens_per_minute_per_workspace`.

## Retrieve one model

```bash theme={"dark"}
curl https://api.runinfra.ai/v1/models/nemotron-3-5-lightning-30b \
  -H "Authorization: Bearer $RUNINFRA_GATEWAY_KEY"
```

Returns a single object with the same fields and no `list` wrapper. An id your key cannot reach returns `404` `model_not_found`. URL-encode any id containing a slash.

## Authentication and rate limits

The endpoint needs a workspace API key, sent as `Authorization: Bearer` or `x-api-key`. It uses that key's requests-per-minute limit and returns `X-RateLimit-Limit`, `X-RateLimit-Remaining`, `X-RateLimit-Reset` and `X-RateLimit-Tier`. Listing models reserves no inference credit, but it is still an authenticated request against your per-minute budget, so do not call it in a hot loop.

## Related

<Columns cols={3}>
  <Card title="Chat completions" icon="braces" href="/docs/api-reference/chat-completions">
    Send one of these ids as the `model` field.
  </Card>

  <Card title="Authentication" icon="key" href="/docs/api-reference/authentication">
    The key that decides which ids you see.
  </Card>

  <Card title="Rate limits" icon="gauge" href="/docs/api-reference/rate-limits">
    The four layers that can refuse you.
  </Card>
</Columns>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.