> ## Documentation Index
> Fetch the complete documentation index at: https://runinfra.ai/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Chat completions

> The request contract for POST /v1/chat/completions on hosted Model APIs.

```http theme={"dark"}
POST https://api.runinfra.ai/v1/chat/completions
```

This is the endpoint every Model APIs caller reaches. Send a standard chat completion from the SDK you already use, with one of the live model ids from [`GET /v1/models`](/docs/api-reference/models).

<CodeGroup>
  ```python Python theme={"dark"}
  import os
  from openai import OpenAI

  client = OpenAI(
      base_url="https://api.runinfra.ai/v1",
      api_key=os.environ["RUNINFRA_GATEWAY_KEY"],
  )
  response = client.chat.completions.create(
      model="nemotron-3-5-lightning-30b",
      messages=[{"role": "user", "content": "Explain exponential backoff."}],
      max_tokens=16384,
  )
  print(response.choices[0].message.content)
  ```

  ```typescript TypeScript theme={"dark"}
  import OpenAI from "openai";

  const client = new OpenAI({
    baseURL: "https://api.runinfra.ai/v1",
    apiKey: process.env.RUNINFRA_GATEWAY_KEY,
  });
  const response = await client.chat.completions.create({
    model: "nemotron-3-5-lightning-30b",
    messages: [{ role: "user", content: "Explain exponential backoff." }],
    max_tokens: 16384,
  });
  console.log(response.choices[0]?.message?.content);
  ```

  ```bash cURL theme={"dark"}
  curl https://api.runinfra.ai/v1/chat/completions \
    -H "Authorization: Bearer $RUNINFRA_GATEWAY_KEY" \
    -H "Content-Type: application/json" \
    -d '{"model":"nemotron-3-5-lightning-30b","messages":[{"role":"user","content":"Explain exponential backoff."}],"max_tokens":16384}'
  ```
</CodeGroup>

## What each model supports

The model-specific behavior and dated measurements below also cover models that may be paused. A capability does not imply current availability.

<Note>
  Run `GET /v1/models` for the current model list and each model's capabilities. A model that is temporarily paused is reported there and on its [Model Library](https://runinfra.ai/inference-api) page.
</Note>

Requesting a capability a model does not have returns `400` with `error.code` `hosted_capability_not_supported` and `param` naming the field. You find out at the call, not in the output.

## Give reasoning models room

Reasoning models can spend their output budget before they answer. Nemotron 3.5 Lightning 30B reasons by default. Reasoning tokens count toward `max_tokens` before the answer, so set `max_tokens` to at least the figure in the table below. A smaller budget can be spent entirely on reasoning and return a completion whose `content` is empty.

<div className="block dark:hidden">
  <svg viewBox="0 0 720 218" width="100%" role="img" aria-label="Two output budgets for a reasoning model: a 2,048-token budget can be fully consumed by reasoning, while a 16,384-token budget leaves more room for the answer. The diagram illustrates measurements from August 14, 2026, when the probes billed output. Under the current usage contract, a response with no answer settles at zero." fill="none" xmlns="http://www.w3.org/2000/svg"><text x="24" y="16" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">Reasoning models: output budget</text><text x="696" y="16" fill="#78786f" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0" textAnchor="end">reasoning tokens count toward max\_tokens</text><text x="24" y="51" fill="#78786f" fontFamily="Consolas, Menlo, monospace" fontSize="10">max\_tokens: 2048</text><text x="696" y="51" fill="#78786f" fontFamily="Consolas, Menlo, monospace" fontSize="10" textAnchor="end">budget exhausted</text><rect x="24" y="58" width="672" height="4" fill="#b07f24" shapeRendering="crispEdges" /><rect x="24" y="74" width="5" height="5" fill="#b95f5f" shapeRendering="crispEdges" /><text x="36" y="82" fill="#52524c" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">The whole budget went to reasoning: no answer. No-answer responses now settle at zero.</text><text x="24" y="119" fill="#78786f" fontFamily="Consolas, Menlo, monospace" fontSize="10">max\_tokens: 16384</text><text x="696" y="119" fill="#78786f" fontFamily="Consolas, Menlo, monospace" fontSize="10" textAnchor="end">answer delivered</text><rect x="24" y="126" width="202" height="4" fill="#b07f24" shapeRendering="crispEdges" /><rect x="226" y="126" width="370" height="4" fill="#76b900" shapeRendering="crispEdges" /><rect x="596" y="126" width="100" height="4" fill="#efefe9" shapeRendering="crispEdges" /><rect x="24" y="142" width="5" height="5" fill="#b07f24" shapeRendering="crispEdges" /><text x="36" y="150" fill="#52524c" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">reasoning</text><rect x="116" y="142" width="5" height="5" fill="#76b900" shapeRendering="crispEdges" /><text x="128" y="150" fill="#52524c" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">answer</text><text x="196" y="150" fill="#78786f" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">Each reasoning model's page publishes its recommended minimum.</text><line x1="24" y1="172" x2="696" y2="172" stroke="#e8e8e3" strokeWidth="1" strokeDasharray="3 3" /><rect x="21.5" y="169.5" width="5" height="5" fill="#bbb9b1" shapeRendering="crispEdges" /><rect x="693.5" y="169.5" width="5" height="5" fill="#bbb9b1" shapeRendering="crispEdges" /><text x="24" y="190" fill="#6e6d64" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">Measured 2026-08-14: at 2,048 tokens both public models returned an empty or truncated answer.</text><text x="24" y="204" fill="#6e6d64" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">Those probes billed output. DeepSeek V4 Flash: use 16,384. Segment widths are illustrative.</text></svg>
</div>

<div className="hidden dark:block">
  <svg viewBox="0 0 720 218" width="100%" role="img" aria-label="Two output budgets for a reasoning model: a 2,048-token budget can be fully consumed by reasoning, while a 16,384-token budget leaves more room for the answer. The diagram illustrates measurements from August 14, 2026, when the probes billed output. Under the current usage contract, a response with no answer settles at zero." fill="none" xmlns="http://www.w3.org/2000/svg"><text x="24" y="16" fill="#6e6d64" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">Reasoning models: output budget</text><text x="696" y="16" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0" textAnchor="end">reasoning tokens count toward max\_tokens</text><text x="24" y="51" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="10">max\_tokens: 2048</text><text x="696" y="51" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="10" textAnchor="end">budget exhausted</text><rect x="24" y="58" width="672" height="4" fill="#d9a64a" shapeRendering="crispEdges" /><rect x="24" y="74" width="5" height="5" fill="#c76f6f" shapeRendering="crispEdges" /><text x="36" y="82" fill="#c8c7ba" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">The whole budget went to reasoning: no answer. No-answer responses now settle at zero.</text><text x="24" y="119" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="10">max\_tokens: 16384</text><text x="696" y="119" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="10" textAnchor="end">answer delivered</text><rect x="24" y="126" width="202" height="4" fill="#d9a64a" shapeRendering="crispEdges" /><rect x="226" y="126" width="370" height="4" fill="#76b900" shapeRendering="crispEdges" /><rect x="596" y="126" width="100" height="4" fill="#2b2b28" shapeRendering="crispEdges" /><rect x="24" y="142" width="5" height="5" fill="#d9a64a" shapeRendering="crispEdges" /><text x="36" y="150" fill="#c8c7ba" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">reasoning</text><rect x="116" y="142" width="5" height="5" fill="#76b900" shapeRendering="crispEdges" /><text x="128" y="150" fill="#c8c7ba" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">answer</text><text x="196" y="150" fill="#9a998e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">Each reasoning model's page publishes its recommended minimum.</text><line x1="24" y1="172" x2="696" y2="172" stroke="#383833" strokeWidth="1" strokeDasharray="3 3" /><rect x="21.5" y="169.5" width="5" height="5" fill="#6e6d64" shapeRendering="crispEdges" /><rect x="693.5" y="169.5" width="5" height="5" fill="#6e6d64" shapeRendering="crispEdges" /><text x="24" y="190" fill="#9a998e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">Measured 2026-08-14: at 2,048 tokens both public models returned an empty or truncated answer.</text><text x="24" y="204" fill="#9a998e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">Those probes billed output. DeepSeek V4 Flash: use 16,384. Segment widths are illustrative.</text></svg>
</div>

Measured on this endpoint on 2026-08-14 by sending the prompt "Write a TypeScript retry helper." and recording how often a complete answer came back:

| Model | Use at least | What smaller budgets did |
| - | - | - |
| DeepSeek V4 Flash | **16,384** | At 2,048 and 4,096 the answer was empty every time. At 8,192 it was empty half the time. |
| `nemotron-3-5-lightning-30b` | **4,096** | At 2,048 the answer was empty in 2 of 5 runs and cut off mid-sentence in 4 of 5. |

For another model, read the recommended output budget on its page in the [Model Library](https://runinfra.ai/inference-api).

The budget is a ceiling, not a spend. At 16,384 DeepSeek V4 Flash still stopped on its own after roughly 7,500 output tokens, and raising the ceiling further did not make it generate more. A ceiling below the table's figure produces the reasoning and not the answer.

<Accordion title="How to detect an exhausted budget in code">
  RunInfra preserves `finish_reason: "length"` and adds `choices[].runinfra.output_status` when a choice reaches its generation limit before producing final answer content.

  ```json theme={"dark"}
  {
    "choices": [
      {
        "finish_reason": "length",
        "message": { "content": null },
        "runinfra": {
          "output_status": {
            "code": "generation_limit_reached_before_answer",
            "message": "The model reached its generation limit before producing answer content. Increase max_tokens or max_completion_tokens if the model limit allows, or shorten the prompt, then retry."
          }
        }
      }
    ],
    "usage": {
      "completion_tokens": 300,
      "runinfra": {
        "output_token_accounting": {
          "visible_answer_tokens": 0,
          "non_answer_completion_tokens": 300,
          "sources": {
            "visible_answer_tokens": "proxy_classified",
            "non_answer_completion_tokens": "provider_reported"
          }
        }
      }
    }
  }
  ```

  Branch on `choices[].runinfra.output_status.code`. `generation_limit_reached_before_answer` means the choice ended for `length`. If another terminal `finish_reason` produced no answer content, the code is `no_answer_content` and you should inspect `finish_reason` before retrying.

  The aggregate `usage` annotation appears only when every choice lacks a final answer. RunInfra does not estimate a reasoning-token count: `completion_tokens` stays the model-reported output total, and `non_answer_completion_tokens` records that those tokens produced no answer. The annotation changes no pricing; a response with no billable output settles at zero, as described below.
</Accordion>

## The cost of the request, and its cached input, on its usage

Every hosted model response carries the calculated charge for the request. On credits, it uses the settlement formula over the tokens the provider reported. On a coding plan or Standby, the calculated charge is zero and the usage value is reported separately. A response with no billable output settles at zero and prints `0`. Beside the cost, every hosted response carries the count of input tokens that were billed at the cached input rate.

```json theme={"dark"}
{
  "usage": {
    "prompt_tokens": 20000,
    "completion_tokens": 500,
    "total_tokens": 20500,
    "prompt_tokens_details": {
      "cached_tokens": 0
    },
    "cost": 0.00294,
    "runinfra": {
      "cost_microcents": 294000,
      "cached_input_tokens": 0
    }
  }
}
```

| Field | Type | Value |
| - | - | - |
| `usage.cost` | number | US dollars, to eight decimals. One microcent is exactly 0.00000001 dollars, so nothing is rounded away. |
| `usage.runinfra.cost_microcents` | integer | Calculated charge in microcents: one cent is 1,000,000 microcents. Zero when the plan or Standby paid. |
| `usage.runinfra.plan_value_microcents` | integer | Usage value at the applicable per-token rates. Present only when the plan or Standby paid. |
| `usage.runinfra.paid_from` | string | `plan` or `standby`, present only with plan-funded or Standby-funded usage. Credit-funded usage retains the ordinary cost fields. |
| `usage.prompt_tokens_details.cached_tokens` | integer | Input tokens billed at the cached input rate, in the field OpenAI clients already read. Present on every hosted response; `0` when no input was billed as cached. |
| `usage.runinfra.cached_input_tokens` | integer | The same count, beside the cost, for a reader that only looks at the `runinfra` object. |

The cost covers the tokens the provider reported, at the model's prices for uncached input, cached input (when the model discloses a cache hit) and output. During a promotional free window it is `0`, present, not absent. These cost and cache fields describe hosted models.

The cached count is the number the cost was computed with. On a billable response, `prompt_tokens - cached_tokens` is exactly what was billed at the uncached input rate. The cached count is `0` rather than absent whenever nothing was billed as cached: a first send, a prompt too short to match, a response that settled at zero, or a model whose cache is shared across tenants. On a shared cache a hit is neither disclosed nor priced, so `0` there means no cached discount was applied, not that the cache did nothing. `/v1/models` says which models publish an isolated cache.

`cost_microcents` is the canonical figure and `cost` is the same amount expressed in dollars. Your balance is debited in whole cents, and the fraction of a cent left over carries to your next request, so one debit can differ from one request's cost by less than a cent. Over any period the two agree to within one cent of carry.

On a stream, the cost and the cached count ride on the usage frame that is sent together with `[DONE]`, so send `stream_options: {"include_usage": true}` to receive them (the one usage frame forwarded without it, when a stream produced no answer, carries them too). Content chunks carry neither; a running figure on every delta would be a guess. A stream that ends in an error never quotes a cost. A `/v1/responses` usage carries the cost in the same two fields, and the cached count as `input_tokens_details.cached_tokens` and `runinfra.cached_input_tokens`. The usage of an [idempotent replay](/docs/api-reference/idempotent-retries) of a streamed request (`X-RunInfra-Idempotent-Replay: true`) carries the figures that were settled.

Read the [funding headers](/docs/api-reference/usage#response-funding-headers) for the request's admission decision. [Plan usage](/docs/api-reference/usage) reports current plan limits and policy; [Credits and budget](/docs/api-reference/credits) reports the credit balance.

## Turning reasoning off

`reasoning_effort: "none"` turns reasoning off for one request. The completion then carries the answer and no reasoning stream. Measured on this endpoint on 2026-08-16 with an identical prompt at temperature 0, the completion fell from 35 tokens to 3 on DeepSeek V4 Flash, from 335 to 5 on `nemotron-3-5-lightning-30b`, and from 67 to 5 on `qwen3-8-27b`. On Qwen3.8 Flash Next, measured 2026-08-29 on a one-line prompt, the completion fell from 30 tokens at the default effort to 2 with `none`. Because reasoning bills at the output rate, this is the largest per-request cost lever on short tasks.

Two models cannot turn reasoning off. Qwen3.8 2.4T A95B and GLM 5.3 Flash both refuse `reasoning_effort: "none"` with `400` rather than accepting it and reasoning anyway, so the request is never billed for effort you asked to skip. Both refusals name the smallest budget they accept. Send `"low"` for the smallest reasoning budget on either.

<Accordion title="How the effort levels behave between none and the ceiling">
  Measured on 2026-08-16 at temperature 0 with repeated runs per level, with the later measurements dated below.

  * DeepSeek V4 Flash applies `max` when you send no `reasoning_effort`, and treats `high` and `xhigh` as `max`. An omitted field on a request that can call tools (`tools` sent and `tool_choice` not `"none"`) without a JSON `response_format` runs at `medium`. `medium` is reproducibly distinct from the default.
  * DeepSeek V4 Pro answers without reasoning at the default, so there is nothing to turn off; the levels are not measured.
  * Qwen3.8 2.4T A95B produces reproducibly distinct reasoning at `low`, `medium` and `xhigh` (113, 140 and 89 completion tokens on the probe, with `xhigh` equal to the omitted default). `high` is sent as `xhigh`. `none` is refused.
  * `qwen3-8-27b` measurably alters its reasoning at `medium`, against a twice-identical baseline. `high` and `max` are sent as `xhigh`, the model's maximum and its default.
  * `nemotron-3-5-lightning-30b` level deltas stayed inside the model's own run-to-run variance, so treat effort there as the off switch only.
  * GLM 5.3 Flash: `low` and `high` are the two tiers the model honors as sent. An omitted field and every other accepted value, `medium` included, run at the model's maximum effort. `none` is refused. Send `low` for the smallest budget. Measured on 2026-08-27.
  * `ornith-1-5-35b` reasons at the default and `none` turns it off; the levels between are not measured.
  * Qwen3.8 Flash Next: `none` is accepted and returns the answer directly. `high` and `max` are sent as `xhigh`, the model's maximum and its default. `low` and `medium` are accepted as sent. `minimal` is refused with `400` and `param: "reasoning_effort"`. Measured on 2026-08-30.

  The vendor spellings `enable_thinking`, `chat_template_kwargs.enable_thinking`, `thinking_budget` and `thinking` are accepted for request-shape compatibility and have no effect on any listed model: a request relying on them runs at the model's default effort and can still produce billable reasoning. Send `reasoning_effort` instead; it is the one honored spelling.
</Accordion>

## What the gateway does with a field you send

<div className="block dark:hidden">
  <svg viewBox="0 0 720 246" width="100%" role="img" aria-label="Each top-level request field meets one of three fates: validated by the gateway and forwarded, forwarded untouched from the allowlist, or removed before forwarding because it is not on the allowlist." fill="none" xmlns="http://www.w3.org/2000/svg"><text x="24" y="16" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">What happens to a field you send</text><text x="696" y="16" fill="#78786f" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0" textAnchor="end">POST /v1/chat/completions</text><rect x="24.5" y="34.5" width="215" height="151" fill="#ffffff" stroke="#e8e8e3" strokeWidth="1" shapeRendering="crispEdges" /><rect x="38" y="59" width="5" height="5" fill="#76b900" shapeRendering="crispEdges" /><text x="52" y="68" fill="#0f0f0e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="12.5" fontWeight="500" letterSpacing="-0.12">Checked, then sent</text><text x="38" y="95" fill="#52524c" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">The gateway validates the type</text><text x="38" y="112" fill="#52524c" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">or the range first. A bad value</text><text x="38" y="129" fill="#52524c" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">is refused with 400 before the</text><text x="38" y="146" fill="#52524c" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">request is served or billed.</text><text x="38" y="172" fill="#78786f" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0">model, messages, max\_tokens</text><rect x="252.5" y="34.5" width="215" height="151" fill="#ffffff" stroke="#e8e8e3" strokeWidth="1" shapeRendering="crispEdges" /><rect x="266" y="59" width="5" height="5" fill="#6e6d64" shapeRendering="crispEdges" /><text x="280" y="68" fill="#0f0f0e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="12.5" fontWeight="500" letterSpacing="-0.12">Sent untouched</text><text x="266" y="95" fill="#52524c" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">On the forwarding allowlist.</text><text x="266" y="112" fill="#52524c" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">Any JSON value, no gateway</text><text x="266" y="129" fill="#52524c" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">check. The model can still</text><text x="266" y="146" fill="#52524c" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">reject it, as a 400.</text><text x="266" y="172" fill="#78786f" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0">seed, user, service\_tier</text><rect x="480.5" y="34.5" width="215" height="151" fill="#ffffff" stroke="#e8e8e3" strokeWidth="1" shapeRendering="crispEdges" /><rect x="494" y="59" width="5" height="5" fill="#b07f24" shapeRendering="crispEdges" /><text x="508" y="68" fill="#0f0f0e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="12.5" fontWeight="500" letterSpacing="-0.12">Removed</text><text x="494" y="95" fill="#52524c" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">Not on the allowlist, so it is</text><text x="494" y="112" fill="#52524c" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">dropped before forwarding. The</text><text x="494" y="129" fill="#52524c" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">model never sees it and you</text><text x="494" y="146" fill="#52524c" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">get no warning.</text><text x="494" y="172" fill="#78786f" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0">top\_k, guided\_json, reasoning</text><line x1="24" y1="200" x2="696" y2="200" stroke="#e8e8e3" strokeWidth="1" strokeDasharray="3 3" /><rect x="21.5" y="197.5" width="5" height="5" fill="#bbb9b1" shapeRendering="crispEdges" /><rect x="693.5" y="197.5" width="5" height="5" fill="#bbb9b1" shapeRendering="crispEdges" /><text x="24" y="218" fill="#6e6d64" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">The allowlist is exhaustive. An unknown top-level field passes request validation and is still removed</text><text x="24" y="232" fill="#6e6d64" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">before the request reaches the model. Check the list before depending on a field.</text></svg>
</div>

<div className="hidden dark:block">
  <svg viewBox="0 0 720 246" width="100%" role="img" aria-label="Each top-level request field meets one of three fates: validated by the gateway and forwarded, forwarded untouched from the allowlist, or removed before forwarding because it is not on the allowlist." fill="none" xmlns="http://www.w3.org/2000/svg"><text x="24" y="16" fill="#6e6d64" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">What happens to a field you send</text><text x="696" y="16" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0" textAnchor="end">POST /v1/chat/completions</text><rect x="24.5" y="34.5" width="215" height="151" fill="#161614" stroke="#383833" strokeWidth="1" shapeRendering="crispEdges" /><rect x="38" y="59" width="5" height="5" fill="#76b900" shapeRendering="crispEdges" /><text x="52" y="68" fill="#f0efe2" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="12.5" fontWeight="500" letterSpacing="-0.12">Checked, then sent</text><text x="38" y="95" fill="#c8c7ba" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">The gateway validates the type</text><text x="38" y="112" fill="#c8c7ba" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">or the range first. A bad value</text><text x="38" y="129" fill="#c8c7ba" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">is refused with 400 before the</text><text x="38" y="146" fill="#c8c7ba" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">request is served or billed.</text><text x="38" y="172" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0">model, messages, max\_tokens</text><rect x="252.5" y="34.5" width="215" height="151" fill="#161614" stroke="#383833" strokeWidth="1" shapeRendering="crispEdges" /><rect x="266" y="59" width="5" height="5" fill="#9a998e" shapeRendering="crispEdges" /><text x="280" y="68" fill="#f0efe2" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="12.5" fontWeight="500" letterSpacing="-0.12">Sent untouched</text><text x="266" y="95" fill="#c8c7ba" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">On the forwarding allowlist.</text><text x="266" y="112" fill="#c8c7ba" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">Any JSON value, no gateway</text><text x="266" y="129" fill="#c8c7ba" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">check. The model can still</text><text x="266" y="146" fill="#c8c7ba" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">reject it, as a 400.</text><text x="266" y="172" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0">seed, user, service\_tier</text><rect x="480.5" y="34.5" width="215" height="151" fill="#161614" stroke="#383833" strokeWidth="1" shapeRendering="crispEdges" /><rect x="494" y="59" width="5" height="5" fill="#d9a64a" shapeRendering="crispEdges" /><text x="508" y="68" fill="#f0efe2" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="12.5" fontWeight="500" letterSpacing="-0.12">Removed</text><text x="494" y="95" fill="#c8c7ba" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">Not on the allowlist, so it is</text><text x="494" y="112" fill="#c8c7ba" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">dropped before forwarding. The</text><text x="494" y="129" fill="#c8c7ba" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">model never sees it and you</text><text x="494" y="146" fill="#c8c7ba" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">get no warning.</text><text x="494" y="172" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0">top\_k, guided\_json, reasoning</text><line x1="24" y1="200" x2="696" y2="200" stroke="#383833" strokeWidth="1" strokeDasharray="3 3" /><rect x="21.5" y="197.5" width="5" height="5" fill="#6e6d64" shapeRendering="crispEdges" /><rect x="693.5" y="197.5" width="5" height="5" fill="#6e6d64" shapeRendering="crispEdges" /><text x="24" y="218" fill="#9a998e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">The allowlist is exhaustive. An unknown top-level field passes request validation and is still removed</text><text x="24" y="232" fill="#9a998e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">before the request reaches the model. Check the list before depending on a field.</text></svg>
</div>

Every field lands in one of three dispositions. These are checked against the contract below, then forwarded:

| Field | Type | Default | Behavior |
| - | - | - | - |
| `model` | string | none | Required. Must resolve to a model your key can reach. |
| `messages` | object array | none | Required, at least one message. Extra message fields pass through. Some models need system messages in a set order, see [System message placement](#system-message-placement). |
| `stream` | boolean | `false` | `true` returns an SSE response. |
| `temperature` | number, 0 to 2 | none | Range checked, then forwarded. |
| `top_p` | number, 0 to 1 | none | Range checked, then forwarded. |
| `max_tokens` | integer, 1 or more | none, see [Output limits](#output-limits) | Per-generation output limit, counts toward the total output budget. |
| `max_completion_tokens` | integer, 1 or more | none, see [Output limits](#output-limits) | Same. If you send both, the larger is used for the budget check. |
| `n` | integer, 1 to 4 | `1` | Generated choices. The total output budget covers all of them. |
| `best_of` | integer, 1 to 4 | `1` | Forwarded when present, included in the choice budget. |
| `stop` | string or string array | none | At most 4 sequences, 256 characters each. Beyond either bound returns `400`. |

A message's `role` is `system`, `developer`, `user`, `assistant` or `tool`. Any other role returns `400`. `content` may be a string, an array, or `null`. An assistant message needs `content` or a non-empty `tool_calls`. Each `tool_calls[].function.arguments` must hold a JSON object, such as `"{}"` for a call with no arguments. A tool message needs `content` and a non-empty string `tool_call_id`. Every other role needs `content`.

### System message placement

A `developer` message is sent to the model as a `system` message. Most models accept system messages anywhere in the conversation and receive them where you put them.

Qwen3.8 27B, Qwen3.8 2.4T A95B, Qwen3.8 Flash Next, and Ornith 1.5 35B accept one system message, and only as the first message. For these models the gateway repairs the order before the request reaches the model:

* Consecutive system messages at the start are merged into one, in order, separated by a blank line.
* A system message later in the conversation is sent as a `user` message in the same position.

Every other message, tool call, and tool result keeps its position and content. A repaired conversation still extends the previous turn's prompt, so prompt caching keeps working across turns. Coding agents that send more than one system message, such as Claude Code and Codex CLI, run on these models without extra configuration. A leading system message that contains anything other than text is not merged, and the request returns `400` with `param: "messages"`.

<AccordionGroup>
  <Accordion title="Sent untouched: the forwarding allowlist">
    The gateway accepts any JSON value on these, applies no default, and forwards it unchanged. The model can still reject the value, which comes back as `400 invalid_request_error`.

    `audio`, `frequency_penalty`, `function_call`, `functions`, `logit_bias`, `logprobs`, `metadata`, `modalities`, `moderation`, `parallel_tool_calls`, `prediction`, `presence_penalty`, `safety_identifier`, `seed`, `service_tier`, `store`, `tool_choice`, `top_logprobs`, `user`, `verbosity`, `web_search_options`.

    Some allowlisted fields do carry gateway behavior:

    | Field | Behavior |
    | - | - |
    | `functions`, `function_call` | Forwarded unchanged, except on DeepSeek V4 Flash, which refuses them with `400 hosted_parameter_not_supported` unless a non-empty `tools` array is also sent. Send `tools` and `tool_choice` instead. |
    | `logit_bias` | Forwarded unchanged, except on DeepSeek V4 Flash and Qwen3.8 Flash Next, where a non-empty `logit_bias` is refused with `400 hosted_parameter_not_supported`: those two models do not support per-token bias. An empty object is forwarded. |
    | `logprobs`, `top_logprobs` | Forwarded unchanged, except on `ornith-1-5-35b` and DeepSeek V4 Pro, which refuse a logprobs request with `400 hosted_parameter_not_supported`. |
    | `reasoning_effort` | Forwarded unchanged, except where a model publishes a default or a rewrite. On DeepSeek V4 Flash, omitting it applies `max`. A request that can call tools (`tools` sent and `tool_choice` not `"none"`) without a JSON `response_format` gets `medium` instead, reported in `X-RunInfra-Reasoning-Effort-Applied`. `high`, `xhigh` and `max` all go upstream as `max`. |
    | `response_format` | Forwarded to a model that can enforce it. If a model needs a conforming effort for `json_schema` or `json_object`, omitting `reasoning_effort` applies `"none"` on DeepSeek V4 Flash and `nemotron-3-5-lightning-30b`, or `"low"` on GLM 5.3 Flash, and reports it in `X-RunInfra-Reasoning-Effort-Applied`. GLM also accepts an explicit `"high"`. An incompatible explicit effort is refused with `400`. |
    | `stream_options` | The gateway always requests usage so it can settle. Client-visible usage normally needs `include_usage: true`; a stream with no final answer also forwards an annotated usage frame without opt-in. |
    | `tool_choice` | Forwarded unchanged, except `"required"` on `nemotron-3-5-lightning-30b`, which is refused with `400 hosted_parameter_not_supported`. Send `"auto"` or name one function. |
    | `tools` | An empty array is refused with `400 hosted_parameter_not_supported`. Anything else is forwarded unchanged, with two exceptions. A bound pair in a tool's parameter schema that no value can satisfy (`minimum` above `maximum` or an exclusive bound that closes the range, `minItems` above `maxItems`, `minLength` above `maxLength`, `minProperties` above `maxProperties`) is removed before the request is forwarded, because a forced tool call enforces the schema during generation and a range no value can satisfy cannot be enforced; satisfiable bounds are kept. An `enum` with no values or a `pattern` that is not a valid regular expression cannot be repaired and is refused with `400 hosted_parameter_not_supported` naming the path. Both rules apply to a `json_schema` response format as well. Forwarding does not guarantee the model accepts a particular tool definition. |
  </Accordion>

  <Accordion title="Removed before the request leaves the gateway">
    The allowlist is exhaustive: every other top-level field is dropped. Common examples are `max_output_tokens`, `reasoning`, `text`, `include`, `truncation` and `previous_response_id` (Responses API fields on the Chat route); `top_k`, `min_p`, `repetition_penalty` and `length_penalty`; any nonstandard guided-decoding or beam-search field such as `guided_json`, `guided_regex`, `guided_choice`, `guided_grammar` and `use_beam_search` (use `response_format` and `tools` for constrained output); and `num_return_sequences`, `num_beams` and `beam_width`, which the gateway reads for the four-choice safety check and then removes.

    An unknown field can pass request validation and still be removed here, so never rely on silent pass-through for a field that is not on the allowlist.
  </Accordion>
</AccordionGroup>

## Output limits

| Limit | Value |
| - | - |
| Total output tokens per request | **The model's remaining context**, that is its context window minus your prompt, across all generated choices. Oversized values are clamped, never refused. |
| Generated choices per request | **4** |
| Full response | **740 seconds**, then `504` `hosted_inference_response_timeout` |
| First token on a stream | **180 seconds**, then `504` `hosted_inference_response_timeout` |
| JSON request body | **3.5MB**, then `413` `payload_too_large` |
| Images per request | **8**, counted across every message in the request, then `400` `hosted_parameter_not_supported` |
| `stop` sequences | 4, of 256 characters each |

The body-size and image limits count every message in the request, including the images and history a client resends each turn. A single image fits at about 2.6MB on disk because base64 adds a third; drop older images from the conversation or split them across requests.

Omit both output-token fields and the gateway forwards no output limit at all, so the model generates until it stops on its own or reaches the remaining context. Send an oversized `max_tokens` and the gateway clamps it to what the context allows rather than refusing the request: the response carries `X-RunInfra-Output-Clamped: true` and `X-RunInfra-Output-Token-Ceiling` with the value used.

Two bounds apply at once and the smaller one wins. Tokens are bounded by the remaining context above; wall-clock is bounded by the 740 second response limit, which at typical decode rates is the binding limit for very long generations. `/v1/models` publishes both, `max_output_tokens` beside `response_time_ceiling_seconds`, which carries the same 740 second figure, so size a long generation against the pair rather than the token figure alone.

The `504` message names the budget that was exceeded. For a non-streaming request, the error message suggests two remedies: `stream: true` (response headers arrive with the first token) and a smaller `max_tokens`. A first-token timeout on a stream suggests a smaller prompt, or retrying when the model has more capacity.

A model's **context window** is a separate per-model limit, published on that model's page in the [Model Library](https://runinfra.ai/inference-api). Served context is a property of how the model is served, not of the weights, so read it on the model page rather than a model card.

## Three things a request can be refused for

Each one is refused up front with a `400` naming the field, rather than answered with output you cannot use.

<AccordionGroup>
  <Accordion title="A JSON format the model cannot hold while reasoning">
    `response_format` of `json_schema` or `json_object` is enforced during generation, and on some models that enforcement cannot start until the model has finished reasoning. On those models, omitting `reasoning_effort` applies an effort measured to return a conforming object: `reasoning_effort: "none"` on DeepSeek V4 Flash and `nemotron-3-5-lightning-30b`, and `"low"` on GLM 5.3 Flash, where an explicit `"high"` also conforms. The gateway reports an automatically applied value in the `X-RunInfra-Reasoning-Effort-Applied` response header. Send an incompatible effort beside the format and the request is refused with `400`, `code` `hosted_parameter_not_supported` and `param: "response_format"`.

    The fix is the effort value sent alongside `response_format`:

    ```json theme={"dark"}
    {
      "model": "nemotron-3-5-lightning-30b",
      "messages": [{ "role": "user", "content": "Classify and rewrite this ticket." }],
      "reasoning_effort": "none",
      "response_format": {
        "type": "json_schema",
        "json_schema": {
          "name": "triage",
          "strict": true,
          "schema": {
            "type": "object",
            "properties": {
              "category": { "type": "integer", "enum": [1, 2, 3, 4] },
              "rewrite": { "type": "string" }
            },
            "required": ["category", "rewrite"],
            "additionalProperties": false
          }
        }
      }
    }
    ```

    Check each model's JSON mode requirements in the [Model Library](https://runinfra.ai/inference-api) before changing the model in this request. `response_format: {"type": "text"}` constrains nothing and is never refused. On `/v1/responses` the same remedy is a top-level `reasoning_effort` beside `text.format`. Working snippets are in [Structured output](/docs/cookbook/structured-output).
  </Accordion>

  <Accordion title="An image, audio, or video content part">
    Accepted input is a per-model fact, stated on each model page as the Accepted input row. Qwen3.8 27B, Ornith 1.5 35B, GLM 5.3 Flash and Qwen3.8 Flash Next accept `image_url` and `input_image` content parts as inline data URLs, up to 8 images per request: `{"type": "image_url", "image_url": {"url": "data:image/png;base64,..."}}` with png, jpeg, webp or gif. Remote image URLs are not fetched; download the image, base64-encode it, and send the data URL. Image tokens are counted by the server inside `prompt_tokens` and bill at the model's input rate. Each image must be a complete file: a png cut inside one of its data chunks, a jpeg without its end-of-image marker, or a webp shorter than its header declares is refused with a 400 that names the message and part (for example `messages[2].content[1]`) before any generation starts, and an image that still cannot be decoded is refused the same way rather than reported as unavailable.

    On every other live model, and for audio and video everywhere, a non-text content part inside `messages` returns `400` with `param: "messages"` and `code: "hosted_parameter_not_supported"`, naming the model and the part to remove. The part is refused rather than stripped, so a request carrying media is never billed as if it were text.
  </Accordion>

  <Accordion title="A prompt-cache control field">
    `prompt_cache_options` and `prompt_cache_retention` are rejected with `400` `hosted_parameter_not_supported`. `prompt_cache_key` is accepted: when no [session header](/docs/api-reference/session-affinity) supplies an identity, the gateway reads it as a cache-affinity hint and removes it before the request reaches the model. It grants no control over retention. Prefix caching, where a model has it, is automatic server behavior and is not controllable per request.
  </Accordion>
</AccordionGroup>

## Prefix cache retention

Retention is what `GET /v1/models` states as `cache_tiers` and `cache_retention`. DeepSeek V4 Pro reports `cache_retention: "tiered_host_memory"`: an idle session's cached prefix survives beyond immediate memory pressure and can be reused when the session resumes. Every other chat model, GLM 5.3 Flash and Qwen3.8 Flash Next included, reports `"best_effort"`: cached prefixes are retained opportunistically and may be evicted at any time, with no retention promise. Keeping a session's follow-up turns close to its warm cache is covered in [Session affinity](/docs/api-reference/session-affinity); read `usage.prompt_tokens_details.cached_tokens` for what the cache did on each request.

## Calling a paused model

A hosted model can be temporarily paused. The request returns `503` with `code` `hosted_model_paused` and a `Retry-After` header carrying the seconds until availability is next rechecked, capped at 3,600, or 60 when no recheck time is known or the known one has already passed. `error.paused_until` names that next recheck time, present only when one is scheduled; it is a recheck deadline, not a promised return. Nothing is charged. The model id stays valid and stays listed by [`GET /v1/models`](/docs/api-reference/models) with `availability: "paused"`, so retry rather than re-resolving your configuration.

## Related

<Columns cols={3}>
  <Card title="Streaming" icon="radio" href="/docs/api-reference/streaming">
    Read deltas and ask for the usage frame.
  </Card>

  <Card title="Rate limits" icon="gauge" href="/docs/api-reference/rate-limits">
    The four layers that can refuse you, and how to back off.
  </Card>

  <Card title="Errors" icon="circle-alert" href="/docs/api-reference/errors">
    Every status, code, and caller action.
  </Card>
</Columns>
