> ## Documentation Index
> Fetch the complete documentation index at: https://runinfra.ai/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Data retention

> Hosted Model APIs are zero data retention by default. What we store per request, the windows you open, and what we never store.

**Hosted Model APIs are zero data retention by default.** Your prompt text is never stored. We keep per-request usage metadata, which is what billing and support need. We do not train on your requests.

For a security review, this is everything that outlives a request by default:

* **Usage metadata and keyed prompt digests**, kept until you delete your account. The digests cannot be reversed into your text.
* **Cached prompt prefixes**, held in the serving GPU's memory between your requests, isolated to your workspace, until evicted. They are never shared across workspaces.

Content outlives a request only in windows you open: the model's output (a non-streamed response body, up to 6 MiB) for up to 24 hours when you send an `Idempotency-Key`, and audio uploaded through the large-file lane for up to 2 hours (no hosted model accepts uploads today).

## What we store per request

One usage record per request, plus the billing transaction that settles it. Between them they hold the fields below and settlement metadata: where the token counts came from, how a stream ended, interim token counts, and the settlement timestamp. None of these values contains your content.

| What | Why it exists |
| - | - |
| Token counts | Input and output tokens are what you are billed on. On models that publish a cached rate, the cached input count is stored when the model reported one. |
| Cost | The amount actually charged, and the per-million-token rates it was charged at. |
| Timing | Request latency, so you can see performance in your usage view. |
| Model called | Which model served the request. |
| Request id | The value returned in your `x-request-id` header, so a specific call can be found if you dispute a charge or open a ticket. |
| Status | The HTTP status the request finished with, and whether it settled. |
| Key and workspace | Which API key and workspace to attribute the spend to. Ids only. |

Both records also carry free-form metadata. It holds identifiers, billing bookkeeping (token counts, the price snapshot, settlement state), and machine-readable outcome codes. No prompt, completion, or header is written to these records. The one text-shaped value is a redacted, length-capped copy of the serving side's own rejection sentence when a request is refused as invalid; it describes the refusal, never your content. The groups below describe these values.

| Metadata group | What it holds |
| - | - |
| Pre-serve hold markers | The request id, the credit held while the request runs, a flag recording whether a promotional price applied and, when the call is exempt from a hold or is billed on input only, a marker saying so. Written before the request is served. |
| Serving evidence | Timing measurements such as time to first token, cache-related token counts, the reason generation stopped, how a stream ended, error and outcome codes, pseudonymous diagnostics for investigating the request, and keyed, workspace-scoped digests of the prompt's messages, system prompt and tools, which cannot confirm a guessed prompt without our key and are kept until you delete your account. None contains your text. Each value is present only when it was observed. |
| Rejection detail | Where a rejection originated and, for a rejection from the model, the sanitized reason sentence, capped in length. That sentence is wording our own code produced, not a copy of your request. |
| Replay markers | The identity of the original request and a flag marking a duplicate, so a retried call is recognized. |

The pseudonymous identifier cannot be reversed to recover anything you sent, but we still treat it as pseudonymous data, never as anonymous.

The billing transaction carries the same settled counts, price snapshot, status, and pseudonymous identifier. Neither record receives your messages, your files, or the model's output.

## The replay window

Sending an `Idempotency-Key` holds that request's response for up to 24 hours, so a retry with the same key returns the same answer instead of running, and charging for, the work twice. Send no key and no response body is held.

* **It holds the response, not the request.** The request body itself is not kept; only a fingerprint of it that cannot be turned back into the body is, so a key reused with different parameters can be rejected rather than answered with the wrong reply.
* **It expires on its own.** An entry is gone 24 hours after it is written.
* **Streamed responses are never stored.** The cache refuses an event stream outright. A duplicate of a streamed request receives a receipt instead: the usage and the cost of the original, with the message content explicitly `null`, because we did not keep the generated text and will not invent it.
* **Responses larger than 6 MiB are not stored.** Only the byte count is kept, and a duplicate is told the original is too large to replay.
* **It is scoped to your workspace and your key.** An entry is only ever visible to the workspace and key that wrote it, so no other tenant can reach it.
* **It is purpose bound.** The only thing that reads an entry is the code path answering a duplicate of the same request. It feeds no analytics, no training, and no support tooling.
* **What it holds is the response we built, not the one the model returned.** Every response is projected through an allowlist before it leaves us: only the documented chat completion response fields survive, plus our own `runinfra` namespace. Any field beyond that allowlist, including any echo of your prompt that arrives with the model output, is dropped by construction rather than by a rule someone remembered to write, so it is neither in your reply nor in the cache entry.

## The audio upload window

Audio uploaded through the large-file lane is held for up to 2 hours after the upload, and a successful transcription deletes it immediately. No hosted model serves audio transcription today, so the lane accepts no uploads. Audio sent as a direct `file` part exists only for that request.

## What we never store

| What | Status |
| - | - |
| Prompt text, including system and tool messages | Never stored |
| Images you send | Never stored |
| Audio sent as a direct `file` part | Never stored |
| Streamed completion text, including reasoning and tool call arguments | Never stored |
| Non-streamed completion text, when you send no `Idempotency-Key` | Never stored |
| Your content in analytics or product telemetry | Never collected |
| Your content in crash and error reporting | Never collected |

Audio uploaded through the large-file lane is the one input with a window of its own, described under [The audio upload window](#the-audio-upload-window).

Error reporting is configured so that a server exception cannot carry request bodies or local variables with it, so request content cannot leak into it as a side effect. Product analytics carries no user-written text at all.

## What exists only while the request runs

Serving a request means holding it in memory for as long as the request takes. That is unavoidable, and it is the boundary of any inference provider's retention claim.

Your messages exist only on the serving side for as long as the call takes and are released with it. Repeated leading context is reused as a **cached prefix**, held only on the serving side and never written to our database, and it is what makes cached input cheaper than fresh input.

On every model where we publish a **cached input price**, cached prefixes in the `gpu` tier are isolated per workspace, so a prefix in that tier is only ever reused by the workspace that created it. The cached token count and the cached rate are enabled only while that isolation is active. A published cached rate signals this isolation; any additional cache tier's scope is stated under [Cached working state](#cached-working-state). On models with no cached rate, no cached input discount is billed and no cache hit is disclosed; the billed cached-token count is `0`.

<span id="on-the-serving-machine" />

## Cached working state

The serving side does not receive your workspace id. Where it needs a per-workspace value to keep cached prefixes separate, it receives a value derived from your workspace id that cannot be reversed to it.

Being precise about what that buys you: **the value identifies no one to the serving side, and no one to anybody outside RunInfra.** It is not anonymous to us. We can reproduce it for a workspace id we already know and match it back. That is deliberate, because it is what lets us investigate a cache isolation question, and it is why we call the value unlinkable by the recipient rather than unlinkable full stop.

DeepSeek V4 Pro keeps its prefix cache across its `gpu` and `host_ram` tiers between requests, evicted under capacity pressure, with no keep window. Per-workspace isolation on this model covers the `gpu` tier; the `host_ram` tier that extends it is not isolated per workspace today.

On every other model with a published cached input rate, including GLM 5.3 Flash and Qwen3.8 Flash Next, the listing reports `cache_isolation` as `isolated`; the cached prefix uses the published `gpu` tier and is evicted under capacity pressure rather than kept for a window. A model with no cached rate is treated as shared: cache hits are not disclosed or discounted, and the billed cached-token count is `0`.

[`GET /v1/models`](/docs/api-reference/models) states each chat model's behavior as `cache_isolation`, `cache_tiers`, and `cache_retention`; `cache_isolation` reads `shared` whenever no cached rate is published. `cache_tiers` and `cache_retention` carry what has been proven for each model's prefix cache; DeepSeek V4 Pro, for example, reports the `gpu` and `host_ram` tiers.

## Rejected and failed requests

**A rejected request writes no text from your request.** When a request is refused as invalid with `400` or `422`, what reaches our operational log is the request id, the status, the error type, the error code, your workspace id, the model id you asked for, and the name of the field or capability the error points to. When the refusal is raised after the request was dispatched, because a message or image could not be accepted or the model rejected the request, the message you receive is logged with them, after a credential redactor and a length cap. That message is wording our own code produces, not a copy of your request.

On a `5xx` we log the error object our code or the upstream produced, passed through a credential redactor first. That is diagnostic text about the failure, not your request: no code path puts your messages into it.

## Scope

* Retention terms in writing, for a procurement or compliance review, are a contract conversation. [Contact us](https://runinfra.ai/contact).
* Zero data retention by default means no request or response content is stored unless you open one of the windows above. Billing metadata is always kept, and it holds no content.
* Your own copies are yours to manage. If you log requests and responses on your side, that retention is governed by your systems.

## Related

<Columns cols={2}>
  <Card title="Idempotent retries" icon="repeat" href="/docs/api-reference/idempotent-retries">
    The header that opens the 24 hour window, and what it replays.
  </Card>

  <Card title="Authentication" icon="key" href="/docs/api-reference/authentication">
    How keys are stored, rotated, and retired.
  </Card>
</Columns>
