> ## Documentation Index
> Fetch the complete documentation index at: https://runinfra.ai/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Session affinity

> Keep a session's requests together so its prefix cache stays warm across turns and fan-out workers.

Send an `x-session-id` header with a stable value of your choosing, and every request carrying that value is kept together as one session. On models with prefix caching, this helps keep the session's cache warm across turns, which can lower time to first token and cached input cost on multi-turn conversations and fan-out agent workloads. The header is optional, it works on streaming and non-streaming requests alike, and there is nothing to create or clean up. A session exists by being used.

The same header applies to [`/v1/messages`](/docs/api-reference/anthropic-messages). Send `x-session-id` there too: `metadata.user_id` reaches the model as `user` and is not a routing identity on its own.

Claude Code, Codex CLI, Grok Build, Mistral Vibe, and OpenCode already send session headers of their own, and the gateway reads them, so their sessions stay together with no configuration. See [Coding agents](#coding-agents).

<CodeGroup>
  ```python Python theme={"dark"}
  response = client.chat.completions.create(
      model="nemotron-3-5-lightning-30b",
      messages=[{"role": "user", "content": "Summarize the thread so far."}],
      max_tokens=16384,
      extra_headers={"x-session-id": "chat-7c2f1a"},
  )
  ```

  ```typescript TypeScript theme={"dark"}
  const response = await client.chat.completions.create(
    {
      model: "nemotron-3-5-lightning-30b",
      messages: [{ role: "user", content: "Summarize the thread so far." }],
      max_tokens: 16384,
    },
    { headers: { "x-session-id": "chat-7c2f1a" } },
  );
  ```

  ```bash cURL theme={"dark"}
  curl https://api.runinfra.ai/v1/chat/completions \
    -H "Authorization: Bearer $RUNINFRA_GATEWAY_KEY" \
    -H "Content-Type: application/json" \
    -H "x-session-id: chat-7c2f1a" \
    -d '{"model":"nemotron-3-5-lightning-30b","messages":[{"role":"user","content":"Summarize the thread so far."}],"max_tokens":16384}'
  ```
</CodeGroup>

## The identity headers

The gateway chooses the first valid identity in this order:

| Identity, highest precedence first | What it names |
| - | - |
| `runinfra.sessionId` in the request body | An explicit session identity. |
| `x-runinfra-session` | An explicit session identity, read after the body field. |
| `x-session-id` | This request's own session. Requests sharing a value share the session's warm cache. |
| `x-claude-code-session-id` | Claude Code's conversation. A valid `x-claude-code-agent-id` adds a subagent identity within that conversation. |
| `x-grok-session-id` | Grok Build's session. An empty value means no session. |
| `x-session-affinity` | This request's own session, a spelling OpenCode sends. Prefer `x-session-id` in new code. |
| `x-affinity` | The session header Mistral Vibe sends. |
| `thread-id` | Codex CLI's own conversation thread. Read only alongside a nonempty `x-codex-window-id` header. |
| `session-id` | Codex CLI's fallback session identity. Read only alongside a nonempty `x-codex-window-id` header. |
| `x-parent-session-id` | The parent session this request should co-locate with. Read only when the request names no session of its own, so a fan-out worker reuses its parent's warm prefix. |

The gateway validates each candidate separately. An invalid header is skipped and never masks a valid header lower in the order. It does not cause an error response. After these identities comes the body field `prompt_cache_key`. With none of them, the request may be grouped by its opening messages, described below; otherwise it keeps workspace-level grouping. An identity sharpens that grouping from the workspace to the session.

| Spec | Value |
| - | - |
| Session header length, except `x-runinfra-session` | The value must be nonempty after surrounding whitespace is trimmed. The resolved identity must fit within 256 characters, including an agent namespace where used. |
| Value outside that range, or containing a control character | Ignored. The next identity in the order is read instead, and `prompt_cache_key` and the placement rules below still apply |
| Value shape, `runinfra.sessionId` and `x-runinfra-session` | 1 to 128 characters from letters, digits, `.`, `_`, `:` and `-`. An invalid value is ignored and the session headers are read instead |
| Comparison | Exact and case sensitive |
| No identity sent | May be grouped by the opening messages, otherwise workspace-level grouping |

Claude Code, Grok Build, and Codex identities use separate namespaces so they cannot collide with your own `x-session-id`. Raw session ids never leave the gateway: it hashes each identity with the workspace id before storing routing state or writing logs. Session headers are not forwarded to the model. Recorded agent usage contains header names, counts, and keyed fingerprints, never raw ids, raw header values, or prompt text.

## Fan-out: workers follow the parent

An agent that fans out sends the orchestrator's turns on its own session, and each worker names that session as its parent. The workers then read the same warm prefix the orchestrator built, the shared system prompt and plan, with no coordination between them.

```bash theme={"dark"}
# Orchestrator turn, on its own session
curl https://api.runinfra.ai/v1/chat/completions \
  -H "Authorization: Bearer $RUNINFRA_GATEWAY_KEY" \
  -H "Content-Type: application/json" \
  -H "x-session-id: run-4127" \
  -d '{"model":"nemotron-3-5-lightning-30b","messages":[{"role":"user","content":"Plan the review, then dispatch workers."}],"max_tokens":16384}'

# Worker request, co-located with the parent
curl https://api.runinfra.ai/v1/chat/completions \
  -H "Authorization: Bearer $RUNINFRA_GATEWAY_KEY" \
  -H "Content-Type: application/json" \
  -H "x-parent-session-id: run-4127" \
  -d '{"model":"nemotron-3-5-lightning-30b","messages":[{"role":"user","content":"Review module A against the shared plan."}],"max_tokens":16384}'
```

A request's own `x-session-id` or `x-session-affinity` wins over `x-parent-session-id`, so a worker that should follow its parent sends `x-parent-session-id` only. Give a worker its own `x-session-id` once it carries its own multi-turn context worth keeping warm.

## Coding agents

Claude Code, Codex CLI, Grok Build, Mistral Vibe, and OpenCode send session headers of their own, and the gateway reads them. A session from any of them stays together across its turns with nothing to configure.

* **Claude Code** sends `x-claude-code-session-id` on every request. The main conversation is one session, and each subagent is a session of its own, told apart by the `x-claude-code-agent-id` it adds. A resumed conversation keeps its session.
* **Codex CLI** sends its own conversation thread as `thread-id`. A subagent's `session-id` names the root conversation, so the gateway prefers `thread-id` to keep each subagent's session separate. Both headers require a nonempty `x-codex-window-id`.
* **Grok Build** sends `x-grok-session-id` on every request. When Grok Build has no session, it sends an empty value, which the gateway skips.
* **Mistral Vibe** sends `x-affinity` with its session id on completion and streaming requests.
* **OpenCode** sends its own session as `x-session-id` and `x-session-affinity`, and a subagent adds `x-parent-session-id`. Each subagent keeps its own session.

An `x-session-id` you add yourself wins over an agent's own headers.

## `prompt_cache_key`, `user`, and the opening messages

Clients that already send the standard `prompt_cache_key` request field get the same routing with no new headers: when no session identity is present, `prompt_cache_key` is read as the identity. The field is consumed at the gateway. It shapes routing and is removed before the request reaches the model. `prompt_cache_options` and `prompt_cache_retention` are rejected with `400` `hosted_parameter_not_supported`.

`user` is not a routing identity. It remains a standard request field and reaches the model unchanged. Without a session identity or `prompt_cache_key`, the gateway may place the request by its opening messages instead: the `user` value together with the leading system or developer text and the first user message. Two requests that share a `user` value but open with different messages may be placed apart. A request with no message text keeps workspace-level placement. Prefer `prompt_cache_key`, or a session header, in new code.

A session header keeps an established session's requests close to its warm cache even as load changes, until the session has been idle for 15 minutes or its warm cache becomes unavailable. A body hint such as `prompt_cache_key` groups requests by value but does not hold an established session in place. Neither controls cache retention.

## How placement behaves

* A new session's placement considers current load. An established session keeps its placement while it stays active; load changes alone never move it.
* A session's placement is remembered for 15 minutes of idle time, the window [`GET /v1/models`](/docs/api-reference/models) publishes as `session_affinity_idle_ttl_seconds`. After a longer gap, the next request is placed afresh, with load considered again. This is a routing window, not a cache retention promise; what the cache kept is covered below.
* Affinity is a preference, not a delivery constraint. If the session's warm cache is unavailable, the request is still served and the session settles where it was served, so its cache rebuilds from that request onward.
* Identity values are scoped to your workspace, and the raw value is never stored. The same literal value sent by another workspace shares nothing with yours.

## Response headers

These response headers are not yet live; they will be announced when they ship. Once live, a non-streaming `/v1/chat/completions` response to a request that named a session will report what the prefix cache did for it, on every hosted model whose per-workspace cache isolation is active, the state [`GET /v1/models`](/docs/api-reference/models) reports as `cache_isolation: "isolated"`. Until then, read `usage.prompt_tokens_details.cached_tokens` in the response body, which is live today.

```http theme={"dark"}
x-session-state: <cold | warm | restored>
x-session-tier: <gpu | ram | nvme>
x-session-cached-share: <two-decimal share>
```

| Header | Meaning |
| - | - |
| `x-session-cached-share` | The share of this request's input tokens that were served from cache, to two decimal places, and `0.00` for a request with no input tokens. |
| `x-session-state` | `cold`, `warm` or `restored`. |
| `x-session-tier` | `gpu`, `ram` or `nvme`. The header's `ram` is the tier `GET /v1/models` lists as `host_ram`. |

All three headers are absent, and the response is otherwise unchanged, when any of the following holds:

* The model's per-workspace cache isolation is not active: `cache_isolation` reads `shared` on `GET /v1/models`.
* The request was routed by `prompt_cache_key`, or kept workspace-level placement. A `prompt_cache_key` shapes routing but opens no session, so there is nothing to report. The headers appear when the request named a session, and when it was placed by its opening messages.
* The response is streamed. Streaming responses carry no session headers.
* The call went through `/v1/messages`, which forwards no session headers. They appear only on `/v1/chat/completions` responses.

## Affinity and cache retention

Affinity keeps a session's requests close to the cache that already holds its prefix. Retention is best effort on every model, with no keep window. [`GET /v1/models`](/docs/api-reference/models) states each model's `cache_isolation`, `cache_tiers` and `cache_retention`.

## Related

<Columns cols={2}>
  <Card title="Chat completions" icon="braces" href="/docs/api-reference/chat-completions">
    The request contract these headers ride on.
  </Card>

  <Card title="Idempotent retries" icon="repeat" href="/docs/api-reference/idempotent-retries">
    Retry a call without duplicate work or a second charge.
  </Card>

  <Card title="Prompt caching" icon="database" href="/docs/cookbook/prompt-caching">
    What keeps a prefix reusable across an agent's turns.
  </Card>

  <Card title="Data retention" icon="shield-check" href="/docs/security/data-retention">
    Zero data retention by default, and the windows you open.
  </Card>
</Columns>
