> ## Documentation Index
> Fetch the complete documentation index at: https://runinfra.ai/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Idempotent retries

> Retry a call, streaming or not, without starting duplicate inference work or a second charge.

Send an `Idempotency-Key` header when a lost response or a network failure might make you retry. It is optional, it works on streaming and non-streaming requests alike, and it guarantees one generation and one charge. The same header and replay rules apply to `/v1/messages`.

Use a non-blank printable ASCII value of 255 characters or fewer after trimming. One fresh key per logical request, reused only when you are retrying that same request. A UUID generated at the call site is a convenient source.

<CodeGroup>
  ```python Python theme={"dark"}
  response = client.chat.completions.create(
      model="nemotron-3-5-lightning-30b",
      messages=[{"role": "user", "content": "Write one sentence."}],
      max_tokens=64,
      extra_headers={"Idempotency-Key": "YOUR_STABLE_IDEMPOTENCY_KEY"},
  )
  ```

  ```typescript TypeScript theme={"dark"}
  const response = await client.chat.completions.create(
    {
      model: "nemotron-3-5-lightning-30b",
      messages: [{ role: "user", content: "Write one sentence." }],
      max_tokens: 64,
    },
    { headers: { "Idempotency-Key": "YOUR_STABLE_IDEMPOTENCY_KEY" } },
  );
  ```

  ```bash cURL theme={"dark"}
  curl https://api.runinfra.ai/v1/chat/completions \
    -H "Authorization: Bearer $RUNINFRA_GATEWAY_KEY" \
    -H "Content-Type: application/json" \
    -H "Idempotency-Key: YOUR_STABLE_IDEMPOTENCY_KEY" \
    -d '{"model":"nemotron-3-5-lightning-30b","messages":[{"role":"user","content":"Write one sentence."}],"max_tokens":64}'
  ```
</CodeGroup>

## What a reused key does

<div className="block dark:hidden">
  <svg viewBox="0 0 720 344" width="100%" role="img" aria-label="Idempotent retry outcomes replay a stored response with a replay header, return 409 while the first request is still running, return 400 for different request parameters, and return 422 when the stored response is too large to replay; a streamed request with the same key replays its terminal receipt instead of generating again." fill="none" xmlns="http://www.w3.org/2000/svg"><text x="24" y="16" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">Reuse rules</text><rect x="24.5" y="34.5" width="671" height="47" fill="#ffffff" stroke="#e8e8e3" strokeWidth="1" shapeRendering="crispEdges" /><rect x="38" y="48" width="5" height="5" fill="#76b900" shapeRendering="crispEdges" /><text x="50" y="56" fill="#0f0f0e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="12.5" fontWeight="500" letterSpacing="-0.12">Returns the stored response</text><text x="50" y="74" fill="#5a8f00" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0">X-RunInfra-Idempotent-Replay: true</text><rect x="24.5" y="96.5" width="671" height="47" fill="#ffffff" stroke="#e8e8e3" strokeWidth="1" shapeRendering="crispEdges" /><rect x="38" y="110" width="5" height="5" fill="#b07f24" shapeRendering="crispEdges" /><text x="50" y="118" fill="#b07f24" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0">409 idempotency\_conflict</text><text x="222" y="118" fill="#52524c" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">The first request is still running</text><text x="50" y="136" fill="#78786f" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0">Retry-After: 1</text><text x="150" y="136" fill="#78786f" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">Retry with the same key.</text><rect x="24.5" y="158.5" width="671" height="47" fill="#ffffff" stroke="#e8e8e3" strokeWidth="1" shapeRendering="crispEdges" /><rect x="38" y="172" width="5" height="5" fill="#b95f5f" shapeRendering="crispEdges" /><text x="50" y="180" fill="#b95f5f" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0">400 invalid\_request\_error</text><text x="222" y="180" fill="#52524c" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">The key was used with different request parameters</text><text x="50" y="198" fill="#78786f" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">Use a different key only for a new logical request.</text><rect x="24.5" y="220.5" width="671" height="47" fill="#ffffff" stroke="#e8e8e3" strokeWidth="1" shapeRendering="crispEdges" /><rect x="38" y="234" width="5" height="5" fill="#b95f5f" shapeRendering="crispEdges" /><text x="50" y="242" fill="#b95f5f" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0">422 idempotency\_replay\_unavailable</text><text x="270" y="242" fill="#52524c" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">The stored response is too large to replay</text><text x="50" y="260" fill="#78786f" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">Do not retry with a new key because that can start new inference work.</text><line x1="24" y1="286" x2="696" y2="286" stroke="#e8e8e3" strokeWidth="1" strokeDasharray="3 3" /><rect x="21.5" y="283.5" width="5" height="5" fill="#bbb9b1" shapeRendering="crispEdges" /><rect x="693.5" y="283.5" width="5" height="5" fill="#bbb9b1" shapeRendering="crispEdges" /><rect x="24" y="300" width="5" height="5" fill="#76b900" shapeRendering="crispEdges" /><text x="36" y="308" fill="#52524c" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">A streamed request replays its terminal receipt, never the delivered tokens.</text><text x="36" y="326" fill="#78786f" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0">Idempotency-Key</text><text x="139" y="326" fill="#78786f" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">guarantees one generation and one charge.</text></svg>
</div>

<div className="hidden dark:block">
  <svg viewBox="0 0 720 344" width="100%" role="img" aria-label="Idempotent retry outcomes replay a stored response with a replay header, return 409 while the first request is still running, return 400 for different request parameters, and return 422 when the stored response is too large to replay; a streamed request with the same key replays its terminal receipt instead of generating again." fill="none" xmlns="http://www.w3.org/2000/svg"><text x="24" y="16" fill="#6e6d64" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">Reuse rules</text><rect x="24.5" y="34.5" width="671" height="47" fill="#161614" stroke="#383833" strokeWidth="1" shapeRendering="crispEdges" /><rect x="38" y="48" width="5" height="5" fill="#76b900" shapeRendering="crispEdges" /><text x="50" y="56" fill="#f0efe2" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="12.5" fontWeight="500" letterSpacing="-0.12">Returns the stored response</text><text x="50" y="74" fill="#8fd400" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0">X-RunInfra-Idempotent-Replay: true</text><rect x="24.5" y="96.5" width="671" height="47" fill="#161614" stroke="#383833" strokeWidth="1" shapeRendering="crispEdges" /><rect x="38" y="110" width="5" height="5" fill="#d9a64a" shapeRendering="crispEdges" /><text x="50" y="118" fill="#d9a64a" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0">409 idempotency\_conflict</text><text x="222" y="118" fill="#c8c7ba" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">The first request is still running</text><text x="50" y="136" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0">Retry-After: 1</text><text x="150" y="136" fill="#9a998e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">Retry with the same key.</text><rect x="24.5" y="158.5" width="671" height="47" fill="#161614" stroke="#383833" strokeWidth="1" shapeRendering="crispEdges" /><rect x="38" y="172" width="5" height="5" fill="#c76f6f" shapeRendering="crispEdges" /><text x="50" y="180" fill="#c76f6f" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0">400 invalid\_request\_error</text><text x="222" y="180" fill="#c8c7ba" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">The key was used with different request parameters</text><text x="50" y="198" fill="#9a998e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">Use a different key only for a new logical request.</text><rect x="24.5" y="220.5" width="671" height="47" fill="#161614" stroke="#383833" strokeWidth="1" shapeRendering="crispEdges" /><rect x="38" y="234" width="5" height="5" fill="#c76f6f" shapeRendering="crispEdges" /><text x="50" y="242" fill="#c76f6f" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0">422 idempotency\_replay\_unavailable</text><text x="270" y="242" fill="#c8c7ba" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">The stored response is too large to replay</text><text x="50" y="260" fill="#9a998e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">Do not retry with a new key because that can start new inference work.</text><line x1="24" y1="286" x2="696" y2="286" stroke="#383833" strokeWidth="1" strokeDasharray="3 3" /><rect x="21.5" y="283.5" width="5" height="5" fill="#6e6d64" shapeRendering="crispEdges" /><rect x="693.5" y="283.5" width="5" height="5" fill="#6e6d64" shapeRendering="crispEdges" /><rect x="24" y="300" width="5" height="5" fill="#76b900" shapeRendering="crispEdges" /><text x="36" y="308" fill="#c8c7ba" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">A streamed request replays its terminal receipt, never the delivered tokens.</text><text x="36" y="326" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0">Idempotency-Key</text><text x="139" y="326" fill="#9a998e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">guarantees one generation and one charge.</text></svg>
</div>

Reuse the same key only with the same HTTP method, route, and request body. A key sent with a different body is a different logical request, and it is refused with `400` `idempotency_mismatch` rather than answered with the wrong reply. If you would rather opt out entirely, omit the header: a request without an `Idempotency-Key` runs with no deduplication at all.

## Retrying a stream

A stream is not replayed byte for byte. We do not store the tokens we deliver, so there is nothing to send again, and a live event stream is never stored for replay. What the key protects on a streaming request is the part that costs money.

| Situation | What you get |
| - | - |
| The original stream is still open | `409 idempotency_conflict` with `Retry-After: 1`. No second generation starts. |
| The original stream finished and settled | `200` with `X-RunInfra-Idempotent-Replay: true` and a non-streaming JSON receipt, or a short event stream on a streaming `/v1/messages` retry. Nothing is regenerated and nothing is billed again. |
| The original delivered nothing and was billed nothing | The key is released, so a retry runs normally, exactly as a request without a key would. |

The receipt is an ordinary chat completion envelope, so an OpenAI-compatible client can parse it. `content` is `null` because the delivered text was never stored. `finish_reason` is `length` or `content_filter` when the original reply was cut short, and `null` otherwise. The replay facts ride in `idempotent_replay`.

```json theme={"dark"}
{
  "id": "chatcmpl-8f0a1c2e-...",
  "object": "chat.completion",
  "created": 1786800000,
  "model": "qwen3-8-27b",
  "choices": [
    {
      "index": 0,
      "message": { "role": "assistant", "content": null },
      "finish_reason": null
    }
  ],
  "usage": {
    "prompt_tokens": 1000,
    "completion_tokens": 500,
    "total_tokens": 1500,
    "prompt_tokens_details": { "cached_tokens": 200 },
    "cost": 0.000282,
    "runinfra": { "cost_microcents": 28200, "cached_input_tokens": 200 }
  },
  "idempotent_replay": {
    "replayed": true,
    "streamed": true,
    "content_replayed": false,
    "regenerated": false,
    "billed_again": false,
    "original_status": 200,
    "stream_outcome": "completed",
    "usage_basis": "provider_reported",
    "cost_microcents": 28200,
    "detail": "The original request with this Idempotency-Key streamed to completion and settled once. Streamed tokens are not stored, so this reply carries the terminal usage and cost instead of the generated text. Nothing was regenerated and nothing was billed a second time."
  }
}
```

`X-RunInfra-Idempotent-Replay` is the signal on every route: read it before handing a response to a streaming parser. `/v1/chat/completions` and `/v1/messages` return the `idempotent_replay` object; on `/v1/responses` the reply is converted to a Responses object and only the header survives. A reply the original cut short keeps `status: "incomplete"` and its `incomplete_details.reason` there, as the original stream did. A streaming `/v1/messages` retry receives the receipt as a short event stream with no text.

If you need the generated text after a dropped stream, send a new request with a new key. That is new paid inference. Reusing the old key returns the receipt, not the text.

## Only one charge per key

A retried request is charged once, streamed or not: a duplicate settlement for the same key is refused rather than billed twice. The key holds for the whole time a request can run, so a retry sent while a slow non-streaming request is still generating gets `409 idempotency_conflict` and never starts a second generation.

When a key cannot be checked, a keyed request is refused rather than run without protection. Every endpoint, chat completions, responses and messages included, returns `503 idempotency_unavailable` with `Retry-After` before any inference starts or any charge is made; retry with the same key. On `/v1/messages` the same refusal arrives as Anthropic's `overloaded_error`, which the Anthropic SDKs retry. A request without an `Idempotency-Key` is never affected.

## Client request ids

`X-Client-Request-Id` is an optional correlation header. The gateway echoes it back on the response. It is not a substitute for `Idempotency-Key`. Use a non-blank printable ASCII value of 512 characters or fewer. Any other value is refused with `400` `invalid_trace_header`.

## Related

<Columns cols={2}>
  <Card title="Streaming" icon="radio" href="/docs/api-reference/streaming">
    The frame shapes a retried stream is not replayed as.
  </Card>

  <Card title="Data retention" icon="shield-check" href="/docs/security/data-retention">
    The 24 hour replay window is the one place a response body is stored.
  </Card>
</Columns>
