> ## Documentation Index
> Fetch the complete documentation index at: https://runinfra.ai/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# OpenAI compatibility

> Turn existing OpenAI code into a RunInfra call by changing the base URL, the key, and the model id. Examples for LangChain, LiteLLM, LlamaIndex, the AI SDK, and plain fetch.

If your code already speaks the OpenAI API, three values change and nothing else does.

<CodeGroup>
  ```python Python theme={"dark"}
  import os

  from openai import OpenAI

  client = OpenAI(
      base_url="https://api.runinfra.ai/v1",        # <- change
      api_key=os.environ["RUNINFRA_GATEWAY_KEY"],   # <- change
  )

  response = client.chat.completions.create(
      model="nemotron-3-5-lightning-30b",               # <- change
      messages=[{"role": "user", "content": "Hello"}],
      max_tokens=16384,
  )
  ```

  ```javascript TypeScript theme={"dark"}
  import OpenAI from "openai";

  const client = new OpenAI({
    baseURL: "https://api.runinfra.ai/v1",   // <- change
    apiKey: process.env.RUNINFRA_GATEWAY_KEY, // <- change
  });

  const response = await client.chat.completions.create({
    model: "nemotron-3-5-lightning-30b",              // <- change
    messages: [{ role: "user", content: "Hello" }],
    max_tokens: 16384,
  });
  ```

  ```bash curl theme={"dark"}
  curl https://api.runinfra.ai/v1/chat/completions \
    -H "Authorization: Bearer $RUNINFRA_GATEWAY_KEY" \
    -H "Content-Type: application/json" \
    -d '{"model":"nemotron-3-5-lightning-30b","messages":[{"role":"user","content":"Hello"}],"max_tokens":16384}'
  ```
</CodeGroup>

OpenAI model names do not alias to RunInfra models. Pass a model id from `GET /v1/models`. Streaming, tools, and structured output use the same request and response shapes where the selected model's page lists that capability. Streams terminate with `data: [DONE]`, tool calls appear on the assistant message, and usage appears on the response.

## Set an output budget

This is the one setting that decides whether a reasoning model has room to answer. Reasoning tokens are billed output that count toward `max_tokens`.

<div className="block dark:hidden">
  <svg viewBox="0 0 720 218" width="100%" role="img" aria-label="Two output budgets for a reasoning model: a 2,048-token budget can be fully consumed by reasoning, while a 16,384-token budget leaves more room for the answer. The diagram illustrates measurements from August 14, 2026, when the probes billed output. Under the current usage contract, a response with no answer settles at zero." fill="none" xmlns="http://www.w3.org/2000/svg"><text x="24" y="16" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">Reasoning models: output budget</text><text x="696" y="16" fill="#78786f" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0" textAnchor="end">reasoning tokens count toward max\_tokens</text><text x="24" y="51" fill="#78786f" fontFamily="Consolas, Menlo, monospace" fontSize="10">max\_tokens: 2048</text><text x="696" y="51" fill="#78786f" fontFamily="Consolas, Menlo, monospace" fontSize="10" textAnchor="end">budget exhausted</text><rect x="24" y="58" width="672" height="4" fill="#b07f24" shapeRendering="crispEdges" /><rect x="24" y="74" width="5" height="5" fill="#b95f5f" shapeRendering="crispEdges" /><text x="36" y="82" fill="#52524c" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">The whole budget went to reasoning: no answer. No-answer responses now settle at zero.</text><text x="24" y="119" fill="#78786f" fontFamily="Consolas, Menlo, monospace" fontSize="10">max\_tokens: 16384</text><text x="696" y="119" fill="#78786f" fontFamily="Consolas, Menlo, monospace" fontSize="10" textAnchor="end">answer delivered</text><rect x="24" y="126" width="202" height="4" fill="#b07f24" shapeRendering="crispEdges" /><rect x="226" y="126" width="370" height="4" fill="#76b900" shapeRendering="crispEdges" /><rect x="596" y="126" width="100" height="4" fill="#efefe9" shapeRendering="crispEdges" /><rect x="24" y="142" width="5" height="5" fill="#b07f24" shapeRendering="crispEdges" /><text x="36" y="150" fill="#52524c" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">reasoning</text><rect x="116" y="142" width="5" height="5" fill="#76b900" shapeRendering="crispEdges" /><text x="128" y="150" fill="#52524c" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">answer</text><text x="196" y="150" fill="#78786f" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">Each reasoning model's page publishes its recommended minimum.</text><line x1="24" y1="172" x2="696" y2="172" stroke="#e8e8e3" strokeWidth="1" strokeDasharray="3 3" /><rect x="21.5" y="169.5" width="5" height="5" fill="#bbb9b1" shapeRendering="crispEdges" /><rect x="693.5" y="169.5" width="5" height="5" fill="#bbb9b1" shapeRendering="crispEdges" /><text x="24" y="190" fill="#6e6d64" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">Measured 2026-08-14: at 2,048 tokens both public models returned an empty or truncated answer.</text><text x="24" y="204" fill="#6e6d64" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">Those probes billed output. DeepSeek V4 Flash: use 16,384. Segment widths are illustrative.</text></svg>
</div>

<div className="hidden dark:block">
  <svg viewBox="0 0 720 218" width="100%" role="img" aria-label="Two output budgets for a reasoning model: a 2,048-token budget can be fully consumed by reasoning, while a 16,384-token budget leaves more room for the answer. The diagram illustrates measurements from August 14, 2026, when the probes billed output. Under the current usage contract, a response with no answer settles at zero." fill="none" xmlns="http://www.w3.org/2000/svg"><text x="24" y="16" fill="#6e6d64" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">Reasoning models: output budget</text><text x="696" y="16" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0" textAnchor="end">reasoning tokens count toward max\_tokens</text><text x="24" y="51" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="10">max\_tokens: 2048</text><text x="696" y="51" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="10" textAnchor="end">budget exhausted</text><rect x="24" y="58" width="672" height="4" fill="#d9a64a" shapeRendering="crispEdges" /><rect x="24" y="74" width="5" height="5" fill="#c76f6f" shapeRendering="crispEdges" /><text x="36" y="82" fill="#c8c7ba" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">The whole budget went to reasoning: no answer. No-answer responses now settle at zero.</text><text x="24" y="119" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="10">max\_tokens: 16384</text><text x="696" y="119" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="10" textAnchor="end">answer delivered</text><rect x="24" y="126" width="202" height="4" fill="#d9a64a" shapeRendering="crispEdges" /><rect x="226" y="126" width="370" height="4" fill="#76b900" shapeRendering="crispEdges" /><rect x="596" y="126" width="100" height="4" fill="#2b2b28" shapeRendering="crispEdges" /><rect x="24" y="142" width="5" height="5" fill="#d9a64a" shapeRendering="crispEdges" /><text x="36" y="150" fill="#c8c7ba" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">reasoning</text><rect x="116" y="142" width="5" height="5" fill="#76b900" shapeRendering="crispEdges" /><text x="128" y="150" fill="#c8c7ba" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">answer</text><text x="196" y="150" fill="#9a998e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">Each reasoning model's page publishes its recommended minimum.</text><line x1="24" y1="172" x2="696" y2="172" stroke="#383833" strokeWidth="1" strokeDasharray="3 3" /><rect x="21.5" y="169.5" width="5" height="5" fill="#6e6d64" shapeRendering="crispEdges" /><rect x="693.5" y="169.5" width="5" height="5" fill="#6e6d64" shapeRendering="crispEdges" /><text x="24" y="190" fill="#9a998e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">Measured 2026-08-14: at 2,048 tokens both public models returned an empty or truncated answer.</text><text x="24" y="204" fill="#9a998e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">Those probes billed output. DeepSeek V4 Flash: use 16,384. Segment widths are illustrative.</text></svg>
</div>

For `nemotron-3-5-lightning-30b`, start with at least 4,096 output tokens.

Where a recommended minimum is published, the model's page in the [Model Library](https://runinfra.ai/inference-api) carries it. You pay for tokens generated, not for the ceiling you set, so a generous budget costs nothing extra. If the answer comes back empty, check `finish_reason`: `"length"` means raise the budget, not that the model produced nothing.

## Where the model's thinking arrives

RunInfra emits the model's thinking as **`reasoning`** on the message and on each stream delta. Some OpenAI-compatible clients and providers use the name `reasoning_content` instead, so a client that reads only that name sees the answer but not the thinking. Measured per client:

| Client | Visible answer | Model's thinking |
| - | - | - |
| AI SDK (`@ai-sdk/openai-compatible`) | `text` / `text-delta` | `reasoningText`, and `reasoning-delta` parts |
| OpenAI SDK, Node and Python | `message.content` / `delta.content` | `message.reasoning` / `delta.reasoning` |
| LiteLLM | `message.content` | `message.reasoning_content` |
| LangChain (`langchain-openai`) | `content` | **not exposed**, dropped by the integration |
| Plain `fetch` | `delta.content` | `delta.reasoning` |

The answer is correct in every client above. Only access to the thinking differs, so pick accordingly if you want to render it.

## Prompt caching and session affinity

Prefix caching is automatic server behavior on models with a published cached-input price. You do not select cache placement, retention, or reuse per request: `prompt_cache_options` and `prompt_cache_retention` are rejected with `400` `hosted_parameter_not_supported`.

To maximize hit rate for multi-turn conversations and agent sessions, give repeat requests a consistent session hint so the session's cached prefix is reused:

* Header: `x-session-id: <your-conversation-or-task-id>` (`x-parent-session-id` and the legacy `x-session-affinity` are also accepted)
* Body field: `prompt_cache_key` (the OpenAI SDK spelling), or a stable `user`

All of these work on `/v1/chat/completions` and `/v1/responses`; the hint is used to keep your session's requests and cached prefix together, and it never reaches the model. Without one, repeat requests from the same workspace are still kept together, so steady sessions already get a warm cache; an explicit hint makes that grouping exact for your own conversation or task ids. Cached tokens appear in `usage.prompt_tokens_details.cached_tokens` on `/v1/chat/completions` and in `usage.input_tokens_details.cached_tokens` on `/v1/responses`, and bill at the cached-input rate published on that model's page. The model page and [public pricing page](https://runinfra.ai/pricing) show any measured cache hit-rate figures that are currently published.

## Framework configuration

These framework integrations returned complete answers from `https://api.runinfra.ai/v1` in checks on 2026-08-14. The examples now select models listed as available on September 21, 2026. Versions used: `ai` 7.0.65, `@ai-sdk/openai-compatible` 3.0.30, `openai` 7.4.0 (Node) and 2.54.0 (Python), `langchain-openai`, `litellm`.

### LangChain

```python theme={"dark"}
import os
from langchain_openai import ChatOpenAI

llm = ChatOpenAI(
    model="nemotron-3-5-lightning-30b",
    base_url="https://api.runinfra.ai/v1",
    api_key=os.environ["RUNINFRA_GATEWAY_KEY"],
    max_tokens=4096,
)

print(llm.invoke("Write a TypeScript retry helper.").content)

for chunk in llm.stream("Write a TypeScript retry helper."):
    print(chunk.content, end="", flush=True)
```

### LiteLLM

LiteLLM reaches RunInfra through its `openai/` provider prefix, and exposes the thinking as `reasoning_content`.

```python theme={"dark"}
import os
import litellm

response = litellm.completion(
    model="openai/nemotron-3-5-lightning-30b",
    api_base="https://api.runinfra.ai/v1",
    api_key=os.environ["RUNINFRA_GATEWAY_KEY"],
    messages=[{"role": "user", "content": "Write a TypeScript retry helper."}],
    max_tokens=16384,
)
print(response.choices[0].message.content)
```

Add `stream=True` for deltas, reading `chunk.choices[0].delta.content` and skipping the usage frame, which has no choices.

### LlamaIndex

Use `OpenAILike` from `llama-index-llms-openai-like`, with `is_chat_model=True`. The `OpenAI` class raises `Unknown model` for any id it does not list.

```python theme={"dark"}
import os
from llama_index.llms.openai_like import OpenAILike

llm = OpenAILike(
    model="qwen3-8-27b",
    api_base="https://api.runinfra.ai/v1",
    api_key=os.environ["RUNINFRA_GATEWAY_KEY"],
    is_chat_model=True,
)

print(llm.complete("What is RunInfra?").text)
```

### AI SDK

`@ai-sdk/openai-compatible` reads `reasoning_content` and falls back to `reasoning`, which is the name we emit, so both the answer and the thinking come through.

```typescript theme={"dark"}
// Server-side only: never expose RUNINFRA_GATEWAY_KEY in browser code.
import { createOpenAICompatible } from "@ai-sdk/openai-compatible";
import { streamText } from "ai";

const apiKey = process.env.RUNINFRA_GATEWAY_KEY;
if (!apiKey) throw new Error("Set RUNINFRA_GATEWAY_KEY first.");

const runinfra = createOpenAICompatible({
  name: "runinfra",
  baseURL: "https://api.runinfra.ai/v1",
  apiKey,
});

const result = streamText({
  model: runinfra.chatModel("nemotron-3-5-lightning-30b"),
  maxOutputTokens: 4096,
  prompt: "Write a TypeScript retry helper.",
});

for await (const part of result.fullStream) {
  if (part.type === "reasoning-delta") process.stdout.write(part.text);
  if (part.type === "text-delta") process.stdout.write(part.text);
}
```

`generateText` works the same way, with the thinking on `reasoningText`.

`maxOutputTokens` is the AI SDK's name for the output budget above.

`usage.outputTokenDetails.reasoningTokens` reports `0` even when the model reasoned, because this endpoint does not yet break reasoning out of `completion_tokens`. Total `outputTokens` is correct and is what you are billed for.

### Plain fetch, no SDK

```typescript theme={"dark"}
const response = await fetch("https://api.runinfra.ai/v1/chat/completions", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.RUNINFRA_GATEWAY_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    model: "nemotron-3-5-lightning-30b",
    messages: [{ role: "user", content: "Write a TypeScript retry helper." }],
    max_tokens: 4096,
    stream: true,
    stream_options: { include_usage: true },
  }),
});

if (!response.ok) {
  const { error } = await response.json();
  throw new Error(`${response.status} ${error.code}: ${error.message}`);
}

const reader = response.body.getReader();
const decoder = new TextDecoder();
let buffer = "";

while (true) {
  const { done, value } = await reader.read();
  if (done) break;

  buffer += decoder.decode(value, { stream: true });
  const lines = buffer.split("\n");
  buffer = lines.pop() ?? "";

  for (const line of lines) {
    if (!line.startsWith("data: ")) continue;
    const payload = line.slice(6).trim();
    if (payload === "[DONE]") continue;

    const chunk = JSON.parse(payload);
    const delta = chunk.choices[0]?.delta;
    // The model's thinking arrives as `reasoning`, the answer as `content`.
    if (delta?.reasoning) process.stdout.write(delta.reasoning);
    if (delta?.content) process.stdout.write(delta.content);
  }
}
```

A `data:` line can be split across two network reads, so the trailing partial line stays in `buffer` rather than being parsed. Dropping that carry is the usual cause of intermittent `JSON.parse` failures on a stream that works in testing.

## Where RunInfra differs from OpenAI

<AccordionGroup>
  <Accordion title="Which endpoints exist">
    Model APIs serve `POST /v1/chat/completions`, `POST /v1/responses`, `POST /v1/messages`, `POST /v1/messages/count_tokens`, `GET /v1/models`, `GET /v1/models/{model}`, and [`GET /v1/credits`](/docs/api-reference/credits). Chat completions and Responses are OpenAI-compatible. Messages and token counting use the [Anthropic-compatible contract](/docs/api-reference/anthropic-messages).

    Not supported anywhere: `/v1/completions` (legacy, non-chat), `/v1/files`, `/v1/assistants`, `/v1/threads`, and `/v1/batches`. `/v1/responses` is a chat-completions compatibility adapter and does not implement state, hosted tools (web search, file search), conversation storage, or background jobs.

    On a model whose page lists tool calling, function tools work end to end on `/v1/responses`, in both spellings: declare tools flat (`{"type": "function", "name", "parameters"}`, the Responses shape) or nested under `function` (the chat shape). The model's calls come back as `function_call` output items with `call_id`, `name`, and `arguments`; send results back as `function_call_output` input items with the matching `call_id` (string or content-part array), and resend prior output items in `input` since the endpoint is stateless. Streams carry the full item lifecycle (`response.output_item.added`, `response.content_part.added`, text and `function_call_arguments` deltas, the matching `done` events, and a spec-complete `response.completed`, which ends every stream), so Codex-style clients that build items from `output_item.done` work unchanged. Harness fields: `store` is accepted and always treated as false, nothing is retrievable later; `metadata`, `include` (`reasoning.encrypted_content`, `message.output_text.logprobs`, or empty), `truncation: "disabled"`, `parallel_tool_calls`, `service_tier`, and `client_metadata` are accepted; `truncation: "auto"` and `previous_response_id` are refused with the reason.

    On `/v1/responses`, `output_text` carries the final answer only. The model's thinking arrives as its own `reasoning` output item ahead of the message (raw text under `content` as `reasoning_text` parts, `summary` empty), and on streams as `response.reasoning_text.delta` and `response.reasoning_text.done` events. A completion that spends its whole token budget thinking returns an empty `output_text` with the thinking preserved in the reasoning item, so raise `max_output_tokens` if you see that.

    A reply that `max_output_tokens` or a content filter cut short is not whole, and the response says so: `status` is `"incomplete"` and `incomplete_details.reason` is `max_output_tokens` or `content_filter`. A non-streaming body carries it directly. A stream still ends on `response.completed`, never on `response.incomplete`, and the response inside that event carries the same `status` and reason. Check `status` before you treat an answer as complete. Item status is advisory: only the item being written when the reply stopped is `incomplete`, so parse a function call's `arguments` before you run it.
  </Accordion>

  <Accordion title="What happens to a field you send">
    Fields on the forwarding allowlist reach the model unchanged. Fields that are not on it are dropped before the request leaves the gateway, silently rather than as an error, so a request shape that works today keeps working.

    [Chat completions](/docs/api-reference/chat-completions) lists the allowlist field by field, and names the few that carry gateway behavior. It is the one authoritative list, so this page does not restate it.
  </Accordion>

  <Accordion title="Errors and retries">
    These OpenAI-compatible routes use the OpenAI envelope, `{ error: { message, type, code } }`, with six `error.type` values: `invalid_request_error`, `authentication_error`, `permission_error`, `not_found_error`, `rate_limit_error`, and `api_error`. There is no `server_error` type; 5xx responses carry `api_error`. The Messages routes use the Anthropic envelope. Every status and code is in [Errors](/docs/api-reference/errors), and the triage is in [Troubleshooting](/docs/tips/troubleshooting).

    Retry behavior belongs to your client configuration. Do not blindly retry a charge-bearing request after a partial stream may have reached your app; use an [idempotency key](/docs/api-reference/idempotent-retries) instead. Quote the `X-Request-Id` response header on any support ticket.
  </Accordion>
</AccordionGroup>

<Columns cols={3}>
  <Card title="Hugging Face clients" icon="plug" href="/docs/tools-sdks/hugging-face">
    `huggingface_hub` and `@huggingface/inference`, verified.
  </Card>

  <Card title="Tool calling" icon="wrench" href="/docs/cookbook/tool-calling">
    A complete multi-turn loop.
  </Card>

  <Card title="Limits" icon="gauge" href="/docs/api-reference/rate-limits">
    Context, output, concurrency, tokens per minute.
  </Card>
</Columns>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.