> ## Documentation Index
> Fetch the complete documentation index at: https://runinfra.ai/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Model APIs quickstart

> Make a first hosted chat-completions call in three steps.

Model APIs are open-weight models we host for you. Chat completions and Responses are OpenAI-compatible. Messages is Anthropic-compatible. This quickstart uses `POST /v1/chat/completions` at `https://api.runinfra.ai/v1` with a workspace API key.

<Note>
  Using a coding agent? Your agent can set up RunInfra's open models for you. Copy the prompt from [Set up with your coding agent](/docs/tools-sdks/agent-setup).
</Note>

<Steps>
  <Step title="Get a workspace API key">
    Open the [Model Library](https://runinfra.ai/inference-api), pick a model that is not marked **Paused**, and select **Get API key**. If you already have a key you can copy, the page highlights that key's copy button instead of making another one.
  </Step>

  <Step title="Put the key in your environment">
    ```bash theme={"dark"}
    export RUNINFRA_GATEWAY_KEY="rp_your_key_here"
    ```
  </Step>

  <Step title="Make the call">
    <CodeGroup>
      ```python Python theme={"dark"}
      import os
      from openai import OpenAI

      client = OpenAI(
          base_url="https://api.runinfra.ai/v1",
          api_key=os.environ["RUNINFRA_GATEWAY_KEY"],
      )
      response = client.chat.completions.create(
          model="nemotron-3-5-lightning-30b",
          messages=[{"role": "user", "content": "Write a TypeScript retry helper."}],
          max_tokens=16384,
      )
      print(response.choices[0].message.content)
      ```

      ```typescript TypeScript theme={"dark"}
      import OpenAI from "openai";

      const client = new OpenAI({
        baseURL: "https://api.runinfra.ai/v1",
        apiKey: process.env.RUNINFRA_GATEWAY_KEY,
      });
      const response = await client.chat.completions.create({
        model: "nemotron-3-5-lightning-30b",
        messages: [{ role: "user", content: "Write a TypeScript retry helper." }],
        max_tokens: 16384,
      });
      console.log(response.choices[0]?.message?.content);
      ```

      ```bash cURL theme={"dark"}
      curl https://api.runinfra.ai/v1/chat/completions \
        -H "Authorization: Bearer $RUNINFRA_GATEWAY_KEY" \
        -H "Content-Type: application/json" \
        -d '{"model":"nemotron-3-5-lightning-30b","messages":[{"role":"user","content":"Write a TypeScript retry helper."}],"max_tokens":16384}'
      ```
    </CodeGroup>

    <Accordion title="What comes back">
      A normal completion. Token counts and ids are illustrative; `reasoning` carries the model's thinking, and `content` carries the answer.

      ```json theme={"dark"}
      {
        "id": "chatcmpl-example",
        "object": "chat.completion",
        "model": "nemotron-3-5-lightning-30b",
        "choices": [
          {
            "index": 0,
            "message": {
              "role": "assistant",
              "reasoning": "The user wants a retry helper. I should use exponential backoff...",
              "content": "export async function retry<T>(fn: () => Promise<T>, attempts = 3) { ... }"
            },
            "finish_reason": "stop"
          }
        ],
        "usage": {
          "prompt_tokens": 14,
          "completion_tokens": 892,
          "total_tokens": 906
        }
      }
      ```

      Check `finish_reason` before you use `content`. `stop` means the model finished on its own. `length` means it hit the output budget. On a reasoning model, that can leave `content` empty with the budget spent on thinking, which is why these samples set `max_tokens` to 16,384.
    </Accordion>
  </Step>
</Steps>

## The live chat model ids

Run [`GET /v1/models`](/docs/api-reference/models) for the current list; a model that is temporarily paused is reported there and on its [Model Library](https://runinfra.ai/inference-api) page.

```bash theme={"dark"}
curl https://api.runinfra.ai/v1/models \
  -H "Authorization: Bearer $RUNINFRA_GATEWAY_KEY"
```

## How usage is counted

The token counts in the response `usage` object are the billing counts. For chat completions, cached input is reported at `usage.prompt_tokens_details.cached_tokens` on models that publish a cached-input price, when a cached count is available for the request. Those tokens remain part of `prompt_tokens` and are billed at the model's cached-input rate.

A client-side tokenizer can disagree with the billed count because it may use a different tokenizer revision or omit the exact chat template, special tokens, and media preprocessing that are applied when the request is served. Use the response usage fields and the usage dashboard for reconciliation.

## Why max\_tokens is 16,384

These models reason before they answer, and reasoning tokens are billed output that count toward `max_tokens`.

<div className="block dark:hidden">
  <svg viewBox="0 0 720 218" width="100%" role="img" aria-label="Two output budgets for a reasoning model: a 2,048-token budget can be fully consumed by reasoning, while a 16,384-token budget leaves more room for the answer. The diagram illustrates measurements from August 14, 2026, when the probes billed output. Under the current usage contract, a response with no answer settles at zero." fill="none" xmlns="http://www.w3.org/2000/svg"><text x="24" y="16" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">Reasoning models: output budget</text><text x="696" y="16" fill="#78786f" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0" textAnchor="end">reasoning tokens count toward max\_tokens</text><text x="24" y="51" fill="#78786f" fontFamily="Consolas, Menlo, monospace" fontSize="10">max\_tokens: 2048</text><text x="696" y="51" fill="#78786f" fontFamily="Consolas, Menlo, monospace" fontSize="10" textAnchor="end">budget exhausted</text><rect x="24" y="58" width="672" height="4" fill="#b07f24" shapeRendering="crispEdges" /><rect x="24" y="74" width="5" height="5" fill="#b95f5f" shapeRendering="crispEdges" /><text x="36" y="82" fill="#52524c" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">The whole budget went to reasoning: no answer. No-answer responses now settle at zero.</text><text x="24" y="119" fill="#78786f" fontFamily="Consolas, Menlo, monospace" fontSize="10">max\_tokens: 16384</text><text x="696" y="119" fill="#78786f" fontFamily="Consolas, Menlo, monospace" fontSize="10" textAnchor="end">answer delivered</text><rect x="24" y="126" width="202" height="4" fill="#b07f24" shapeRendering="crispEdges" /><rect x="226" y="126" width="370" height="4" fill="#76b900" shapeRendering="crispEdges" /><rect x="596" y="126" width="100" height="4" fill="#efefe9" shapeRendering="crispEdges" /><rect x="24" y="142" width="5" height="5" fill="#b07f24" shapeRendering="crispEdges" /><text x="36" y="150" fill="#52524c" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">reasoning</text><rect x="116" y="142" width="5" height="5" fill="#76b900" shapeRendering="crispEdges" /><text x="128" y="150" fill="#52524c" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">answer</text><text x="196" y="150" fill="#78786f" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">Each reasoning model's page publishes its recommended minimum.</text><line x1="24" y1="172" x2="696" y2="172" stroke="#e8e8e3" strokeWidth="1" strokeDasharray="3 3" /><rect x="21.5" y="169.5" width="5" height="5" fill="#bbb9b1" shapeRendering="crispEdges" /><rect x="693.5" y="169.5" width="5" height="5" fill="#bbb9b1" shapeRendering="crispEdges" /><text x="24" y="190" fill="#6e6d64" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">Measured 2026-08-14: at 2,048 tokens both public models returned an empty or truncated answer.</text><text x="24" y="204" fill="#6e6d64" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">Those probes billed output. DeepSeek V4 Flash: use 16,384. Segment widths are illustrative.</text></svg>
</div>

<div className="hidden dark:block">
  <svg viewBox="0 0 720 218" width="100%" role="img" aria-label="Two output budgets for a reasoning model: a 2,048-token budget can be fully consumed by reasoning, while a 16,384-token budget leaves more room for the answer. The diagram illustrates measurements from August 14, 2026, when the probes billed output. Under the current usage contract, a response with no answer settles at zero." fill="none" xmlns="http://www.w3.org/2000/svg"><text x="24" y="16" fill="#6e6d64" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">Reasoning models: output budget</text><text x="696" y="16" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0" textAnchor="end">reasoning tokens count toward max\_tokens</text><text x="24" y="51" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="10">max\_tokens: 2048</text><text x="696" y="51" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="10" textAnchor="end">budget exhausted</text><rect x="24" y="58" width="672" height="4" fill="#d9a64a" shapeRendering="crispEdges" /><rect x="24" y="74" width="5" height="5" fill="#c76f6f" shapeRendering="crispEdges" /><text x="36" y="82" fill="#c8c7ba" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">The whole budget went to reasoning: no answer. No-answer responses now settle at zero.</text><text x="24" y="119" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="10">max\_tokens: 16384</text><text x="696" y="119" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="10" textAnchor="end">answer delivered</text><rect x="24" y="126" width="202" height="4" fill="#d9a64a" shapeRendering="crispEdges" /><rect x="226" y="126" width="370" height="4" fill="#76b900" shapeRendering="crispEdges" /><rect x="596" y="126" width="100" height="4" fill="#2b2b28" shapeRendering="crispEdges" /><rect x="24" y="142" width="5" height="5" fill="#d9a64a" shapeRendering="crispEdges" /><text x="36" y="150" fill="#c8c7ba" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">reasoning</text><rect x="116" y="142" width="5" height="5" fill="#76b900" shapeRendering="crispEdges" /><text x="128" y="150" fill="#c8c7ba" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">answer</text><text x="196" y="150" fill="#9a998e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">Each reasoning model's page publishes its recommended minimum.</text><line x1="24" y1="172" x2="696" y2="172" stroke="#383833" strokeWidth="1" strokeDasharray="3 3" /><rect x="21.5" y="169.5" width="5" height="5" fill="#6e6d64" shapeRendering="crispEdges" /><rect x="693.5" y="169.5" width="5" height="5" fill="#6e6d64" shapeRendering="crispEdges" /><text x="24" y="190" fill="#9a998e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">Measured 2026-08-14: at 2,048 tokens both public models returned an empty or truncated answer.</text><text x="24" y="204" fill="#9a998e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">Those probes billed output. DeepSeek V4 Flash: use 16,384. Segment widths are illustrative.</text></svg>
</div>

You pay for what is generated, not for the ceiling you allow, so a generous ceiling costs nothing and a small one can cost you a whole request.

## Next steps

<Columns cols={4}>
  <Card title="Chat completions" icon="messages-square" href="/docs/api-reference/chat-completions">
    Every field the gateway forwards, validates, or drops.
  </Card>

  <Card title="Streaming" icon="radio" href="/docs/api-reference/streaming">
    Read deltas and ask for the usage frame.
  </Card>

  <Card title="Models" icon="list" href="/docs/api-reference/models">
    Discover the ids your key can call.
  </Card>

  <Card title="Coding agents" icon="terminal" href="/docs/tools-sdks/agent-setup">
    Paste one prompt. Your coding agent sets up RunInfra's open models for you.
  </Card>
</Columns>
