> ## Documentation Index
> Fetch the complete documentation index at: https://runinfra.ai/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Streaming

> Stream chat completion deltas over Server-Sent Events, and ask for the usage frame.

For a model whose page lists streaming, set `stream: true` on `POST /v1/chat/completions` to get a Server-Sent Events response. Add `stream_options.include_usage` when your client needs the usage frame. `/v1/messages` uses a different, event-named stream grammar described in [Anthropic Messages](/docs/api-reference/anthropic-messages#streaming).

<CodeGroup>
  ```python Python theme={"dark"}
  import os
  from openai import OpenAI

  client = OpenAI(
      base_url="https://api.runinfra.ai/v1",
      api_key=os.environ["RUNINFRA_GATEWAY_KEY"],
  )
  stream = client.chat.completions.create(
      model="nemotron-3-5-lightning-30b",
      messages=[{"role": "user", "content": "Count from one to three."}],
      max_tokens=64,
      stream=True,
      stream_options={"include_usage": True},
  )

  for chunk in stream:
      if chunk.usage:
          print(f"\nusage: {chunk.usage}")
          continue
      print(chunk.choices[0].delta.content or "", end="", flush=True)
  ```

  ```typescript TypeScript theme={"dark"}
  import OpenAI from "openai";

  const client = new OpenAI({
    baseURL: "https://api.runinfra.ai/v1",
    apiKey: process.env.RUNINFRA_GATEWAY_KEY,
  });
  const stream = await client.chat.completions.create({
    model: "nemotron-3-5-lightning-30b",
    messages: [{ role: "user", content: "Count from one to three." }],
    max_tokens: 64,
    stream: true,
    stream_options: { include_usage: true },
  });

  for await (const chunk of stream) {
    if (chunk.usage) console.error("usage:", chunk.usage);
    process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
  }
  ```

  ```bash cURL theme={"dark"}
  curl -N https://api.runinfra.ai/v1/chat/completions \
    -H "Authorization: Bearer $RUNINFRA_GATEWAY_KEY" \
    -H "Content-Type: application/json" \
    -d '{"model":"nemotron-3-5-lightning-30b","messages":[{"role":"user","content":"Count from one to three."}],"max_tokens":64,"stream":true,"stream_options":{"include_usage":true}}'
  ```
</CodeGroup>

## What arrives, and when

<div className="block dark:hidden">
  <svg viewBox="0 0 720 268" width="100%" role="img" aria-label="The streaming timeline from request to terminal frame: the 180 second first-token budget, the 740 second response budget, the opt-in usage frame with an exception when no final answer was delivered, and the two terminal frames, data DONE or an error frame on the same HTTP 200." fill="none" xmlns="http://www.w3.org/2000/svg"><text x="24" y="16" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">Streaming timeline</text><text x="696" y="16" fill="#78786f" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0" textAnchor="end">stream: true</text><line x1="24" y1="74" x2="696" y2="74" stroke="#e8e8e3" strokeWidth="1" strokeDasharray="3 3" /><rect x="21.5" y="71.5" width="5" height="5" fill="#bbb9b1" shapeRendering="crispEdges" /><rect x="693.5" y="71.5" width="5" height="5" fill="#bbb9b1" shapeRendering="crispEdges" /><rect x="93.5" y="71.5" width="5" height="5" fill="#bbb9b1" shapeRendering="crispEdges" /><line x1="96" y1="77" x2="96" y2="90" stroke="#e8e8e3" strokeWidth="1" /><text x="96" y="104" fill="#0f0f0e" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0" textAnchor="middle">request</text><text x="96" y="118" fill="#78786f" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10" textAnchor="middle">HTTP 200, event stream opens</text><rect x="261.5" y="71.5" width="5" height="5" fill="#bbb9b1" shapeRendering="crispEdges" /><line x1="264" y1="77" x2="264" y2="90" stroke="#e8e8e3" strokeWidth="1" /><text x="264" y="104" fill="#0f0f0e" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0" textAnchor="middle">first token</text><text x="264" y="118" fill="#78786f" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10" textAnchor="middle">within the 180s TTFT budget</text><rect x="429.5" y="71.5" width="5" height="5" fill="#bbb9b1" shapeRendering="crispEdges" /><line x1="432" y1="77" x2="432" y2="90" stroke="#e8e8e3" strokeWidth="1" /><text x="432" y="104" fill="#0f0f0e" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0" textAnchor="middle">delta frames</text><text x="432" y="118" fill="#78786f" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10" textAnchor="middle">one complete frame at a time</text><rect x="573.5" y="71.5" width="5" height="5" fill="#bbb9b1" shapeRendering="crispEdges" /><line x1="576" y1="77" x2="576" y2="90" stroke="#e8e8e3" strokeWidth="1" /><text x="576" y="104" fill="#0f0f0e" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0" textAnchor="middle">usage frame</text><text x="576" y="118" fill="#78786f" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10" textAnchor="middle">opt-in, or no final answer</text><line x1="96" y1="48" x2="664" y2="48" stroke="#bbb9b1" strokeWidth="1" /><line x1="96" y1="44" x2="96" y2="52" stroke="#bbb9b1" strokeWidth="1" /><line x1="664" y1="44" x2="664" y2="52" stroke="#bbb9b1" strokeWidth="1" /><text x="380" y="42" fill="#78786f" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0" textAnchor="middle">740s response budget, then 504</text><text x="24" y="146" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">How a stream ends</text><rect x="24" y="159" width="108" height="3" fill="#76b900" shapeRendering="crispEdges" /><rect x="24" y="162" width="118" height="22" fill="#76b900" shapeRendering="crispEdges" /><rect x="34" y="184" width="108" height="3" fill="#76b900" shapeRendering="crispEdges" /><text x="83" y="177" fill="#0f0f0e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="12" fontWeight="500" textAnchor="middle" letterSpacing="-0.12">data: \[DONE]</text><text x="158" y="177" fill="#52524c" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">the normal terminal frame after the last delta</text><rect x="24" y="210" width="5" height="5" fill="#b95f5f" shapeRendering="crispEdges" /><text x="36" y="218" fill="#b95f5f" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0">error frame</text><text x="126" y="218" fill="#52524c" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">a stream that fails mid-generation ends with one final frame carrying the standard</text><text x="126" y="234" fill="#78786f" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0">upstream\_stream\_error</text><text x="272" y="234" fill="#52524c" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">envelope, on the same HTTP 200. Handle it in your frame loop.</text><text x="24" y="254" fill="#78786f" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">Only tokens actually delivered to you are billed.</text></svg>
</div>

<div className="hidden dark:block">
  <svg viewBox="0 0 720 268" width="100%" role="img" aria-label="The streaming timeline from request to terminal frame: the 180 second first-token budget, the 740 second response budget, the opt-in usage frame with an exception when no final answer was delivered, and the two terminal frames, data DONE or an error frame on the same HTTP 200." fill="none" xmlns="http://www.w3.org/2000/svg"><text x="24" y="16" fill="#6e6d64" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">Streaming timeline</text><text x="696" y="16" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0" textAnchor="end">stream: true</text><line x1="24" y1="74" x2="696" y2="74" stroke="#383833" strokeWidth="1" strokeDasharray="3 3" /><rect x="21.5" y="71.5" width="5" height="5" fill="#6e6d64" shapeRendering="crispEdges" /><rect x="693.5" y="71.5" width="5" height="5" fill="#6e6d64" shapeRendering="crispEdges" /><rect x="93.5" y="71.5" width="5" height="5" fill="#6e6d64" shapeRendering="crispEdges" /><line x1="96" y1="77" x2="96" y2="90" stroke="#383833" strokeWidth="1" /><text x="96" y="104" fill="#f0efe2" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0" textAnchor="middle">request</text><text x="96" y="118" fill="#9a998e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10" textAnchor="middle">HTTP 200, event stream opens</text><rect x="261.5" y="71.5" width="5" height="5" fill="#6e6d64" shapeRendering="crispEdges" /><line x1="264" y1="77" x2="264" y2="90" stroke="#383833" strokeWidth="1" /><text x="264" y="104" fill="#f0efe2" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0" textAnchor="middle">first token</text><text x="264" y="118" fill="#9a998e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10" textAnchor="middle">within the 180s TTFT budget</text><rect x="429.5" y="71.5" width="5" height="5" fill="#6e6d64" shapeRendering="crispEdges" /><line x1="432" y1="77" x2="432" y2="90" stroke="#383833" strokeWidth="1" /><text x="432" y="104" fill="#f0efe2" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0" textAnchor="middle">delta frames</text><text x="432" y="118" fill="#9a998e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10" textAnchor="middle">one complete frame at a time</text><rect x="573.5" y="71.5" width="5" height="5" fill="#6e6d64" shapeRendering="crispEdges" /><line x1="576" y1="77" x2="576" y2="90" stroke="#383833" strokeWidth="1" /><text x="576" y="104" fill="#f0efe2" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0" textAnchor="middle">usage frame</text><text x="576" y="118" fill="#9a998e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10" textAnchor="middle">opt-in, or no final answer</text><line x1="96" y1="48" x2="664" y2="48" stroke="#6e6d64" strokeWidth="1" /><line x1="96" y1="44" x2="96" y2="52" stroke="#6e6d64" strokeWidth="1" /><line x1="664" y1="44" x2="664" y2="52" stroke="#6e6d64" strokeWidth="1" /><text x="380" y="42" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0" textAnchor="middle">740s response budget, then 504</text><text x="24" y="146" fill="#6e6d64" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">How a stream ends</text><rect x="24" y="159" width="108" height="3" fill="#76b900" shapeRendering="crispEdges" /><rect x="24" y="162" width="118" height="22" fill="#76b900" shapeRendering="crispEdges" /><rect x="34" y="184" width="108" height="3" fill="#76b900" shapeRendering="crispEdges" /><text x="83" y="177" fill="#0f0f0e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="12" fontWeight="500" textAnchor="middle" letterSpacing="-0.12">data: \[DONE]</text><text x="158" y="177" fill="#c8c7ba" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">the normal terminal frame after the last delta</text><rect x="24" y="210" width="5" height="5" fill="#c76f6f" shapeRendering="crispEdges" /><text x="36" y="218" fill="#c76f6f" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0">error frame</text><text x="126" y="218" fill="#c8c7ba" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">a stream that fails mid-generation ends with one final frame carrying the standard</text><text x="126" y="234" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0">upstream\_stream\_error</text><text x="272" y="234" fill="#c8c7ba" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">envelope, on the same HTTP 200. Handle it in your frame loop.</text><text x="24" y="254" fill="#9a998e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">Only tokens actually delivered to you are billed.</text></svg>
</div>

## Frame shapes

The response carries `Content-Type: text/event-stream`, `Cache-Control: no-cache` and `Connection: keep-alive`. Content arrives in OpenAI-compatible `data:` frames, and partial output is on `choices[].delta`.

```text theme={"dark"}
data: {"id":"chatcmpl_example","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"One"}}]}
```

The usage frame has an empty `choices` array. On a hosted model it also carries the request's cost and the count of input tokens billed at the cached rate, both described in [Chat completions](/docs/api-reference/chat-completions#the-cost-of-the-request-and-its-cached-input-on-its-usage):

```text theme={"dark"}
data: {"choices":[],"usage":{"prompt_tokens":12,"completion_tokens":4,"total_tokens":16,"prompt_tokens_details":{"cached_tokens":0},"cost":0.0000028,"runinfra":{"cost_microcents":280,"cached_input_tokens":0}}}
```

A completed stream ends with `data: [DONE]`. A stream the gateway terminates early ends with an error frame and no `[DONE]`, so treat a missing `[DONE]` as an incomplete response (see [Errors](/docs/api-reference/errors)).

```text theme={"dark"}
data: [DONE]
```

## Is usage always included?

No. The gateway always asks the model for usage so it can settle the request, but it forwards the usage frame to you only when your request carried `stream_options.include_usage: true`. Omit it and content frames still stream while usage-only frames stay hidden. A usage frame is never synthesized: if the model reports no usage for a request, none is sent.

There is one exception, so you can account for generated tokens even when no answer was delivered. When every choice finishes without final answer content, the gateway forwards one annotated usage frame whether or not you opted in.

## When a stream ends without an answer

A choice that reaches its generation limit before producing answer content carries a machine-readable status on its terminal frame:

```text theme={"dark"}
data: {"choices":[{"index":0,"delta":{},"finish_reason":"length","runinfra":{"output_status":{"code":"generation_limit_reached_before_answer","message":"The model reached its generation limit before producing answer content. Increase max_tokens or max_completion_tokens if the model limit allows, or shorten the prompt, then retry."}}}]}
```

If another terminal `finish_reason` produced no answer, the code is `no_answer_content` instead. Inspect the preserved `finish_reason` before deciding whether to retry.

The annotated usage frame then reports how much of the generated output was not an answer:

```text theme={"dark"}
data: {"choices":[],"usage":{"prompt_tokens":12,"completion_tokens":64,"total_tokens":76,"prompt_tokens_details":{"cached_tokens":0},"cost":0,"runinfra":{"output_token_accounting":{"visible_answer_tokens":0,"non_answer_completion_tokens":64,"sources":{"visible_answer_tokens":"proxy_classified","non_answer_completion_tokens":"provider_reported"}},"cost_microcents":0,"cached_input_tokens":0}}}
```

A stream that produced no answer settles at zero, so this frame says `cost: 0` and `cached_tokens: 0` whatever the model reported: nothing was priced, at either rate.

The usual cause is an output budget too small for a model that reasons first. See [give reasoning models room](/docs/api-reference/chat-completions#give-reasoning-models-room).

<Note>
  An `Idempotency-Key` on a streaming request stops a retry from generating and charging twice, but delivered tokens are never stored and never replayed. Read [Idempotent retries](/docs/api-reference/idempotent-retries) before retrying a stream.
</Note>

## Related

<Columns cols={3}>
  <Card title="Chat completions" icon="braces" href="/docs/api-reference/chat-completions">
    The request contract every stream starts from.
  </Card>

  <Card title="Idempotent retries" icon="refresh-cw" href="/docs/api-reference/idempotent-retries">
    Retry a dropped stream without paying twice.
  </Card>

  <Card title="Errors" icon="circle-alert" href="/docs/api-reference/errors">
    Every status, code, and caller action.
  </Card>
</Columns>
