> ## Documentation Index
> Fetch the complete documentation index at: https://runinfra.ai/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Prompt caching for agents

> Cached input costs about a tenth of uncached on most models. Here is what keeps a prefix reusable across an agent's turns, and what quietly breaks it.

Cached input is billed at the cached rate, roughly a tenth of the standard
input price on most models. For an agent that sends a growing conversation on
every turn, that is the difference between paying full price for your whole
context once and about a tenth of it on every later turn.

Caching is automatic. No header or parameter is required. What
you control is whether your prompt stays *reusable* between turns.

## The rule

The cache matches a **prefix**: the longest run of tokens from the start of
your prompt that is byte-identical to a previous request. Everything from the
first difference onward is recomputed and billed as uncached.

So an agent turn that appends to the end of its context reuses almost
everything. A turn that changes something near the beginning reuses almost
nothing, however small the change.

<div className="block dark:hidden">
  <svg viewBox="0 0 720 358" width="100%" role="img" aria-label="The same agent prompt drawn across three turns on one shared column layout. Turn 1 is the first send, with nothing to match. Turn 2 appends an assistant reply and a second user message, so everything up to the first difference is reused at the cached rate and only the appended tail is recomputed. Turn 2 with a timestamp added at the top of the system block differs at its very first column, so nothing is reused and every token is recomputed. Everything from the first difference onward is recomputed and billed as uncached." fill="none" xmlns="http://www.w3.org/2000/svg"><text x="24" y="16" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">What a prefix cache reuses</text><text x="696" y="16" fill="#78786f" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0" textAnchor="end">the match runs from the start</text><text x="24" y="44" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">Turn 1, first send</text><rect x="24.5" y="52.5" width="111" height="25" fill="#ffffff" stroke="#e8e8e3" strokeWidth="1" shapeRendering="crispEdges" /><text x="80" y="69" fill="#52524c" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" textAnchor="middle">system</text><rect x="136.5" y="52.5" width="95" height="25" fill="#ffffff" stroke="#e8e8e3" strokeWidth="1" shapeRendering="crispEdges" /><text x="184" y="69" fill="#52524c" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" textAnchor="middle">tools</text><rect x="232.5" y="52.5" width="167" height="25" fill="#ffffff" stroke="#e8e8e3" strokeWidth="1" shapeRendering="crispEdges" /><text x="316" y="69" fill="#52524c" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" textAnchor="middle">file context</text><rect x="400.5" y="52.5" width="95" height="25" fill="#ffffff" stroke="#e8e8e3" strokeWidth="1" shapeRendering="crispEdges" /><text x="448" y="69" fill="#52524c" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" textAnchor="middle">user 1</text><line x1="24" y1="92" x2="496" y2="92" stroke="#bbb9b1" strokeWidth="1" /><text x="24" y="106" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">Nothing to match yet</text><text x="24" y="134" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">Turn 2, appended at the end</text><rect x="24.5" y="142.5" width="111" height="25" fill="#ffffff" stroke="#e8e8e3" strokeWidth="1" shapeRendering="crispEdges" /><text x="80" y="159" fill="#52524c" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" textAnchor="middle">system</text><rect x="136.5" y="142.5" width="95" height="25" fill="#ffffff" stroke="#e8e8e3" strokeWidth="1" shapeRendering="crispEdges" /><text x="184" y="159" fill="#52524c" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" textAnchor="middle">tools</text><rect x="232.5" y="142.5" width="167" height="25" fill="#ffffff" stroke="#e8e8e3" strokeWidth="1" shapeRendering="crispEdges" /><text x="316" y="159" fill="#52524c" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" textAnchor="middle">file context</text><rect x="400.5" y="142.5" width="95" height="25" fill="#ffffff" stroke="#e8e8e3" strokeWidth="1" shapeRendering="crispEdges" /><text x="448" y="159" fill="#52524c" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" textAnchor="middle">user 1</text><rect x="496.5" y="142.5" width="95" height="25" fill="#ffffff" stroke="#e8e8e3" strokeWidth="1" shapeRendering="crispEdges" /><text x="544" y="159" fill="#52524c" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" textAnchor="middle">assistant</text><rect x="592.5" y="142.5" width="103" height="25" fill="#ffffff" stroke="#e8e8e3" strokeWidth="1" shapeRendering="crispEdges" /><text x="644" y="159" fill="#52524c" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" textAnchor="middle">user 2</text><line x1="496" y1="138" x2="496" y2="182" stroke="#76b900" strokeWidth="1" /><text x="500" y="134" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">First difference</text><line x1="24" y1="182" x2="496" y2="182" stroke="#76b900" strokeWidth="1" /><text x="24" y="196" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">Reused, cached rate</text><line x1="496" y1="182" x2="696" y2="182" stroke="#b07f24" strokeWidth="1" /><text x="496" y="196" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">Recomputed</text><text x="24" y="224" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">Turn 2, with a timestamp at the top</text><rect x="24.5" y="232.5" width="111" height="25" fill="#ffffff" stroke="#e8e8e3" strokeWidth="1" shapeRendering="crispEdges" /><text x="80" y="249" fill="#52524c" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" textAnchor="middle">system + time</text><rect x="136.5" y="232.5" width="95" height="25" fill="#ffffff" stroke="#e8e8e3" strokeWidth="1" shapeRendering="crispEdges" /><text x="184" y="249" fill="#52524c" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" textAnchor="middle">tools</text><rect x="232.5" y="232.5" width="167" height="25" fill="#ffffff" stroke="#e8e8e3" strokeWidth="1" shapeRendering="crispEdges" /><text x="316" y="249" fill="#52524c" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" textAnchor="middle">file context</text><rect x="400.5" y="232.5" width="95" height="25" fill="#ffffff" stroke="#e8e8e3" strokeWidth="1" shapeRendering="crispEdges" /><text x="448" y="249" fill="#52524c" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" textAnchor="middle">user 1</text><rect x="496.5" y="232.5" width="95" height="25" fill="#ffffff" stroke="#e8e8e3" strokeWidth="1" shapeRendering="crispEdges" /><text x="544" y="249" fill="#52524c" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" textAnchor="middle">assistant</text><rect x="592.5" y="232.5" width="103" height="25" fill="#ffffff" stroke="#e8e8e3" strokeWidth="1" shapeRendering="crispEdges" /><text x="644" y="249" fill="#52524c" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" textAnchor="middle">user 2</text><line x1="24" y1="228" x2="24" y2="272" stroke="#b07f24" strokeWidth="1" /><line x1="24" y1="272" x2="696" y2="272" stroke="#b07f24" strokeWidth="1" /><text x="24" y="286" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">Nothing reused, every token recomputed</text><line x1="24" y1="306" x2="696" y2="306" stroke="#e8e8e3" strokeWidth="1" strokeDasharray="3 3" /><rect x="21.5" y="303.5" width="5" height="5" fill="#bbb9b1" shapeRendering="crispEdges" /><rect x="693.5" y="303.5" width="5" height="5" fill="#bbb9b1" shapeRendering="crispEdges" /><text x="24" y="324" fill="#52524c" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">Everything from the first difference onward is recomputed and billed as uncached.</text><text x="24" y="340" fill="#78786f" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">Column widths are illustrative. Cached input is billed at about a tenth of standard input on most models.</text></svg>
</div>

<div className="hidden dark:block">
  <svg viewBox="0 0 720 358" width="100%" role="img" aria-label="The same agent prompt drawn across three turns on one shared column layout. Turn 1 is the first send, with nothing to match. Turn 2 appends an assistant reply and a second user message, so everything up to the first difference is reused at the cached rate and only the appended tail is recomputed. Turn 2 with a timestamp added at the top of the system block differs at its very first column, so nothing is reused and every token is recomputed. Everything from the first difference onward is recomputed and billed as uncached." fill="none" xmlns="http://www.w3.org/2000/svg"><text x="24" y="16" fill="#6e6d64" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">What a prefix cache reuses</text><text x="696" y="16" fill="#9a998e" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" letterSpacing="0" textAnchor="end">the match runs from the start</text><text x="24" y="44" fill="#6e6d64" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">Turn 1, first send</text><rect x="24.5" y="52.5" width="111" height="25" fill="#161614" stroke="#383833" strokeWidth="1" shapeRendering="crispEdges" /><text x="80" y="69" fill="#c8c7ba" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" textAnchor="middle">system</text><rect x="136.5" y="52.5" width="95" height="25" fill="#161614" stroke="#383833" strokeWidth="1" shapeRendering="crispEdges" /><text x="184" y="69" fill="#c8c7ba" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" textAnchor="middle">tools</text><rect x="232.5" y="52.5" width="167" height="25" fill="#161614" stroke="#383833" strokeWidth="1" shapeRendering="crispEdges" /><text x="316" y="69" fill="#c8c7ba" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" textAnchor="middle">file context</text><rect x="400.5" y="52.5" width="95" height="25" fill="#161614" stroke="#383833" strokeWidth="1" shapeRendering="crispEdges" /><text x="448" y="69" fill="#c8c7ba" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" textAnchor="middle">user 1</text><line x1="24" y1="92" x2="496" y2="92" stroke="#6e6d64" strokeWidth="1" /><text x="24" y="106" fill="#6e6d64" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">Nothing to match yet</text><text x="24" y="134" fill="#6e6d64" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">Turn 2, appended at the end</text><rect x="24.5" y="142.5" width="111" height="25" fill="#161614" stroke="#383833" strokeWidth="1" shapeRendering="crispEdges" /><text x="80" y="159" fill="#c8c7ba" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" textAnchor="middle">system</text><rect x="136.5" y="142.5" width="95" height="25" fill="#161614" stroke="#383833" strokeWidth="1" shapeRendering="crispEdges" /><text x="184" y="159" fill="#c8c7ba" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" textAnchor="middle">tools</text><rect x="232.5" y="142.5" width="167" height="25" fill="#161614" stroke="#383833" strokeWidth="1" shapeRendering="crispEdges" /><text x="316" y="159" fill="#c8c7ba" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" textAnchor="middle">file context</text><rect x="400.5" y="142.5" width="95" height="25" fill="#161614" stroke="#383833" strokeWidth="1" shapeRendering="crispEdges" /><text x="448" y="159" fill="#c8c7ba" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" textAnchor="middle">user 1</text><rect x="496.5" y="142.5" width="95" height="25" fill="#161614" stroke="#383833" strokeWidth="1" shapeRendering="crispEdges" /><text x="544" y="159" fill="#c8c7ba" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" textAnchor="middle">assistant</text><rect x="592.5" y="142.5" width="103" height="25" fill="#161614" stroke="#383833" strokeWidth="1" shapeRendering="crispEdges" /><text x="644" y="159" fill="#c8c7ba" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" textAnchor="middle">user 2</text><line x1="496" y1="138" x2="496" y2="182" stroke="#76b900" strokeWidth="1" /><text x="500" y="134" fill="#6e6d64" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">First difference</text><line x1="24" y1="182" x2="496" y2="182" stroke="#76b900" strokeWidth="1" /><text x="24" y="196" fill="#6e6d64" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">Reused, cached rate</text><line x1="496" y1="182" x2="696" y2="182" stroke="#d9a64a" strokeWidth="1" /><text x="496" y="196" fill="#6e6d64" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">Recomputed</text><text x="24" y="224" fill="#6e6d64" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">Turn 2, with a timestamp at the top</text><rect x="24.5" y="232.5" width="111" height="25" fill="#161614" stroke="#383833" strokeWidth="1" shapeRendering="crispEdges" /><text x="80" y="249" fill="#c8c7ba" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" textAnchor="middle">system + time</text><rect x="136.5" y="232.5" width="95" height="25" fill="#161614" stroke="#383833" strokeWidth="1" shapeRendering="crispEdges" /><text x="184" y="249" fill="#c8c7ba" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" textAnchor="middle">tools</text><rect x="232.5" y="232.5" width="167" height="25" fill="#161614" stroke="#383833" strokeWidth="1" shapeRendering="crispEdges" /><text x="316" y="249" fill="#c8c7ba" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" textAnchor="middle">file context</text><rect x="400.5" y="232.5" width="95" height="25" fill="#161614" stroke="#383833" strokeWidth="1" shapeRendering="crispEdges" /><text x="448" y="249" fill="#c8c7ba" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" textAnchor="middle">user 1</text><rect x="496.5" y="232.5" width="95" height="25" fill="#161614" stroke="#383833" strokeWidth="1" shapeRendering="crispEdges" /><text x="544" y="249" fill="#c8c7ba" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" textAnchor="middle">assistant</text><rect x="592.5" y="232.5" width="103" height="25" fill="#161614" stroke="#383833" strokeWidth="1" shapeRendering="crispEdges" /><text x="644" y="249" fill="#c8c7ba" fontFamily="Consolas, Menlo, monospace" fontSize="10.5" textAnchor="middle">user 2</text><line x1="24" y1="228" x2="24" y2="272" stroke="#d9a64a" strokeWidth="1" /><line x1="24" y1="272" x2="696" y2="272" stroke="#d9a64a" strokeWidth="1" /><text x="24" y="286" fill="#6e6d64" fontFamily="Consolas, Menlo, monospace" fontSize="9" fontWeight="500" letterSpacing="0.3">Nothing reused, every token recomputed</text><line x1="24" y1="306" x2="696" y2="306" stroke="#383833" strokeWidth="1" strokeDasharray="3 3" /><rect x="21.5" y="303.5" width="5" height="5" fill="#6e6d64" shapeRendering="crispEdges" /><rect x="693.5" y="303.5" width="5" height="5" fill="#6e6d64" shapeRendering="crispEdges" /><text x="24" y="324" fill="#c8c7ba" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">Everything from the first difference onward is recomputed and billed as uncached.</text><text x="24" y="340" fill="#9a998e" fontFamily="'Helvetica Neue', Helvetica, Arial, sans-serif" fontSize="10.5" letterSpacing="0">Column widths are illustrative. Cached input is billed at about a tenth of standard input on most models.</text></svg>
</div>

The match is counted in whole cache blocks, and the trailing partial block is
billed at the input rate. The unit size varies by model, so a very short prompt may never hit, and the discount does the most work on long prefixes. Read the cached-token count on your own responses to see how much of your prompt is matching: `usage.prompt_tokens_details.cached_tokens` on Chat Completions, with the Responses and Messages fields under [Reading your own hit rate](#reading-your-own-hit-rate).
One more shape matters on Qwen3.8 Flash Next: a conversation that appends
warms on turn 2, as in the diagram, but a shared prefix re-sent with a
changing suffix, many questions against one document, may take an extra send before `cached_tokens` shows it warm.

## What breaks a prefix

Every item below has been observed reducing a real workload's hit rate. They
are ordered by how often they turn out to be the cause.

<AccordionGroup>
  <Accordion title="A timestamp or date in the system prompt">
    The single most common cause. `Current time: 2026-08-21T09:14:22Z` at the top
    of a system prompt changes on every request, so nothing after it can ever
    match.

    Move it to the **end** of the conversation, as the last user message or a
    trailing system note. It is just as visible to the model there and it stops
    invalidating everything behind it.
  </Accordion>

  <Accordion title="Tool definitions in an unstable order">
    If your tool list is built from a set, a dictionary, or a directory listing,
    its order can change between processes even when the tools do not. Sort tool
    definitions by name before serializing, once, and keep that order for the life
    of the session.
  </Accordion>

  <Accordion title="Non-deterministic JSON serialization">
    `JSON.stringify` over an object preserves insertion order, which can differ
    between runs. Python's `json.dumps` accepts `sort_keys=True`. If you build
    tool schemas or context blocks programmatically, serialize them
    deterministically.

    Whitespace counts too: a pretty-printer that changes indentation between
    versions changes the bytes.
  </Accordion>

  <Accordion title="A session or request id inside the prompt">
    Anything unique per request belongs outside the prompt. If you need the model
    to know a session id, put it at the end, not in the system block.
  </Accordion>

  <Accordion title="Context compaction that rewrites the middle">
    Summarizing older turns is good practice, but replacing the middle of a
    conversation invalidates the prefix from the summary onward. Compact at a
    boundary you keep stable, and prefer appending a summary to rewriting history
    in place.
  </Accordion>

  <Accordion title="Reordering retrieved documents">
    If retrieved context is sorted by a score that shifts slightly between calls,
    the block order changes. Sort retrieved chunks by a stable key, such as
    document id, rather than by score.
  </Accordion>
</AccordionGroup>

## Reading your own hit rate

Chat Completions responses report `usage.prompt_tokens_details.cached_tokens`,
the count of input tokens that request was billed at the cached input rate,
and the same number as `usage.runinfra.cached_input_tokens` beside the cost.
It is `0` rather than absent when nothing was billed as cached. The Responses
API reports it as `usage.input_tokens_details.cached_tokens`, and Anthropic
Messages as `usage.cache_read_input_tokens`. Print it on
every turn: on a workload that should be repetitive, the value that keeps
coming back is where your prompt stops matching itself. The field is
described in [Chat completions](/docs/api-reference/chat-completions#the-cost-of-the-request-and-its-cached-input-on-its-usage).

Your workspace's cache report is on the dashboard's
[Usage Analytics](https://runinfra.ai/inference/usage) page, read with your
signed-in dashboard session rather than an API key. The **Cache by API key**
band carries the workspace totals: cached input tokens, the confirmed hit
rate, confirmed misses, and confirmed savings. The **By API key** and
**By model** tables on the same page carry cached tokens, hit rate, and
savings per key and per model. A low rate on a workload that should be
repetitive means something above is changing.

To compare many requests, record their cached-token counts and look at the
**distribution**, not just an average. Its most useful summary is the mode:
the cached-token count that recurs most often. Because matching is counted in
whole cache blocks, this is roughly where your prompt stops matching itself;
past that point, a field in your prompt may be differing between requests.

<CodeGroup>
  ```python Python theme={"dark"}
  response = client.chat.completions.create(
      model="nemotron-3-5-lightning-30b", messages=messages, max_tokens=16384
  )
  print(response.usage.prompt_tokens_details.cached_tokens)
  ```

  ```typescript TypeScript theme={"dark"}
  const response = await client.chat.completions.create({
    model: "nemotron-3-5-lightning-30b", messages, max_tokens: 16384,
  });
  console.log(response.usage?.prompt_tokens_details?.cached_tokens);
  ```
</CodeGroup>

A worked example from real traffic, measured in August 2026: one workload
showed a mode of **9,984 cached tokens on a fifth of its requests**, against
prompts averaging about 70,000 tokens. Everything up to roughly token 10,000
was reusable and nothing after it was, on every affected turn. That is the
signature of a field changing early in the prompt, not of a cache problem.
A sibling workload on the same model reused 92 percent of a 465,000-token
prompt. These are measured examples, not a promised hit rate.

If the count that recurs sits far below your prompt size, look at what is
between that offset and the start of your context.

## Keeping a session warm

Check current model availability in [`GET /v1/models`](/docs/api-reference/models). The retention behavior below describes configured models and does not guarantee that a model is serving.

On every chat model, retention of a warm prefix is best effort: an idle prefix can be evicted under
capacity pressure, and no lifetime is promised.

On [`/v1/models`](/docs/api-reference/models), `cache_tiers` and `cache_retention`
tell you which kind of cache a model runs. `best_effort` with `cache_tiers`
`["gpu"]` is the GPU copy alone. `tiered_host_memory` means a prefix can outlive the fastest tier: DeepSeek V4 Pro publishes it, and retention there is still best effort.
`session_affinity_idle_ttl_seconds` is a routing window, how long a
quiet session keeps its place close to its cache, not a
retention promise.

The cached-token count on the first turn after a pause is the readout of what
survived it. If the copy was lost while your agent ran tools, waited on a
human, or slept overnight, that turn is billed as a miss; the count and the
bill agree. The matching rule at the top of this page is unchanged: a
returning prompt must be byte-identical to the prefix it wants back.

On any chat model, a stable [`x-session-id`](/docs/api-reference/session-affinity)
keeps a session's follow-up turns close to its warm cache, and pins that placement for the model's published `session_affinity_idle_ttl_seconds` idle window. Without one, a
`prompt_cache_key` in the request body routes the request by its value but
does not pin it. Without either, placement still keeps steady sessions warm;
the full placement order is on the session affinity page.

## What does not affect caching

* **Sampling parameters.** `temperature`, `top_p`, `seed`, and `max_tokens` are
  not part of the prefix.
* **Streaming.** A streamed and a non-streamed request cache identically.
* **Your traffic volume.** Caching is per prefix, not a plan or an allowance.

## Isolation

On models that publish `cache_isolation: "isolated"` in
[`/v1/models`](/docs/api-reference/models), a cached prefix is scoped to your
workspace within the model's published isolation scope. Another customer
does not reuse your isolated prefix, and you do not reuse theirs. That scoping
applies on every request. Read the per-model tier scope in
[Data retention](/docs/security/data-retention#cached-working-state).
On a model that does not publish isolation, a hit is neither disclosed nor
priced. A billable response charges every input token at the standard input
rate. The response reports `cached_tokens: 0`, including when it settles at zero.

## Next steps

<Columns cols={3}>
  <Card title="Session affinity" icon="link" href="/docs/api-reference/session-affinity">
    Keep every turn close to your warm prefix.
  </Card>

  <Card title="Models" icon="list" href="/docs/api-reference/models">
    Read cache isolation, tiers, and retention per model.
  </Card>

  <Card title="Tool calling" icon="wrench" href="/docs/cookbook/tool-calling">
    The agent loop these prefixes come from.
  </Card>
</Columns>
