> ## Documentation Index
> Fetch the complete documentation index at: https://runinfra.ai/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Anthropic Messages

> Send Anthropic-compatible messages, tools, images, thinking controls, and token counts to RunInfra hosted models.

```http theme={"dark"}
POST https://api.runinfra.ai/v1/messages
```

`POST /v1/messages` accepts the Anthropic Messages request and response grammar. It is served by the same models, under the same rules, as `POST /v1/chat/completions`.

Billing, credits, idempotency, rate limits, request-size limits, and cached-input pricing follow the same rules as chat completions. The envelope, content blocks, tool loop, thinking controls, and stream events follow the Anthropic-compatible contract on this page.

<Note>
  This is an Anthropic-compatible API for RunInfra hosted models. It does not claim parity with Anthropic models or every Anthropic platform feature.
</Note>

## Authentication and headers

Send your RunInfra API key in either authentication header on `/v1/messages`, `/v1/messages/count_tokens`, `/v1/models` and `/v1/models/{model}`.

| Header | Behavior |
| - | - |
| `x-api-key: $RUNINFRA_GATEWAY_KEY` | Accepted on those four routes. A non-blank value wins when both authentication headers are present. |
| `Authorization: Bearer $RUNINFRA_GATEWAY_KEY` | Accepted on those four routes and required on every other endpoint. |

The response carries both `request-id` and `x-request-id`. Quote either value when you contact support.

| Compatibility header | Behavior |
| - | - |
| `anthropic-version` | Optional. The response echoes the value you send. It does not select a different grammar. |
| `anthropic-beta` | Accepted and ignored. |
| Claude Code headers | `x-claude-code-session-id` and `x-claude-code-agent-id` route the session, as described in [Coding agents](/docs/api-reference/session-affinity#coding-agents). Other Claude Code headers are accepted and ignored. |

`Idempotency-Key` protects a `/v1/messages` retry from duplicate inference and billing. `x-session-id` uses the same routing-affinity behavior as the OpenAI-compatible routes.

## Create a message

<CodeGroup>
  ```python Python theme={"dark"}
  import os
  from anthropic import Anthropic

  client = Anthropic(
      base_url="https://api.runinfra.ai",
      api_key=os.environ["RUNINFRA_GATEWAY_KEY"],
  )

  message = client.messages.create(
      model="nemotron-3-5-lightning-30b",
      max_tokens=16384,
      system="Answer in one sentence.",
      messages=[
          {"role": "user", "content": "Explain exponential backoff."}
      ],
  )

  for block in message.content:
      if block.type == "text":
          print(block.text)
  ```

  ```typescript TypeScript theme={"dark"}
  import Anthropic from "@anthropic-ai/sdk";

  const client = new Anthropic({
    baseURL: "https://api.runinfra.ai",
    apiKey: process.env.RUNINFRA_GATEWAY_KEY,
  });

  const message = await client.messages.create({
    model: "nemotron-3-5-lightning-30b",
    max_tokens: 16384,
    system: "Answer in one sentence.",
    messages: [
      { role: "user", content: "Explain exponential backoff." },
    ],
  });

  for (const block of message.content) {
    if (block.type === "text") console.log(block.text);
  }
  ```

  ```bash cURL theme={"dark"}
  curl https://api.runinfra.ai/v1/messages \
    -H "x-api-key: $RUNINFRA_GATEWAY_KEY" \
    -H "anthropic-version: 2023-06-01" \
    -H "Content-Type: application/json" \
    -d '{"model":"nemotron-3-5-lightning-30b","max_tokens":16384,"messages":[{"role":"user","content":"Explain exponential backoff."}]}'
  ```
</CodeGroup>

<Note>
  Give the Anthropic SDK the bare host, `https://api.runinfra.ai`. The SDK appends `/v1` itself. OpenAI clients use `https://api.runinfra.ai/v1`.
</Note>

## Request fields

| Field | Required | Behavior |
| - | - | - |
| `model` | yes | A model id available to your key. List ids with [`GET /v1/models`](/docs/api-reference/models). |
| `max_tokens` | yes | An integer of at least `1`. Reasoning spends this budget before the visible answer. `0` is refused. |
| `messages` | yes | At least one `user` or `assistant` message. Consecutive `user` or `assistant` messages are merged into one turn. A `system` message inside `messages` keeps its position; on models that accept a system message only first, the gateway moves it as described in [System message placement](/docs/api-reference/chat-completions#system-message-placement). |
| `system` | no | A string or an array of text blocks. Array blocks remain separate prompt parts. A block-level `cache_control` value is ignored. |
| `tools` | no | Client tool definitions with `name`, optional `description`, and `input_schema`. |
| `tool_choice` | no | `auto`, `none`, `any`, or a named tool. Its optional `disable_parallel_tool_use` field controls whether the model may request several client tools in one turn. |
| `thinking` | no | Enables, adapts, or disables model reasoning as described below. |
| `output_config` | no | Selects a supported reasoning effort or a JSON schema format. |
| `metadata.user_id` | no | A stable routing-affinity hint. |
| `stop_sequences` | no | Passed through as stop strings. The matched value is not reported in the response. |
| `temperature` | no | Passed through. |
| `top_p` | no | Passed through. |
| `top_k` | no | Accepted and ignored. |
| `stream` | no | `true` returns event-named Server-Sent Events. |

Unknown top-level fields are ignored. This keeps newer Anthropic client request shapes usable without forwarding unsupported controls to the model.

## Message content blocks

A message `content` value can be a string or an array of supported blocks.

| Block | Allowed role | Behavior |
| - | - | - |
| `text` | `user`, `assistant` | Passed as text. |
| `image` | `user` | Accepts a base64 source with `image/png`, `image/jpeg`, `image/gif`, or `image/webp` on a model whose page lists image input. |
| `tool_use` | `assistant` | Carries the tool call `id`, `name`, and parsed `input`. |
| `tool_result` | `user` | Carries a matching `tool_use_id`, string content or text and image blocks, and optional `is_error`. Tool results must come before other blocks in that user message. |
| `thinking` | prior `assistant` | Accepted with any `signature`, then dropped before the prior turn reaches the model. |
| `redacted_thinking` | prior `assistant` | Accepted and dropped before the prior turn reaches the model. |

<Note>
  `qwen3-8-27b` and `ornith-1-5-35b` list image input. Check current availability and accepted input on each model's page in the [Model Library](https://runinfra.ai/inference-api) before sending images.
</Note>

An image source with `type: "url"` or `type: "file"` returns `400 invalid_request_error`. A `document` block or another unsupported Anthropic-only block also returns `400` and names the block type.

## Tool use

Declare client tools with JSON Schema input. Fields such as `strict`, `defer_loading`, `cache_control`, and `input_examples` are accepted inside a tool definition and ignored.

```json theme={"dark"}
{
  "model": "nemotron-3-5-lightning-30b",
  "max_tokens": 16384,
  "messages": [
    { "role": "user", "content": "What is the weather in Amman?" }
  ],
  "tools": [
    {
      "name": "get_weather",
      "description": "Read the current weather for a city.",
      "input_schema": {
        "type": "object",
        "properties": {
          "city": { "type": "string" }
        },
        "required": ["city"],
        "additionalProperties": false
      }
    }
  ],
  "tool_choice": { "type": "auto" }
}
```

A tool call returns a `tool_use` content block and `stop_reason: "tool_use"`. Send the result in the next user message before any text or image blocks.

```json theme={"dark"}
{
  "role": "user",
  "content": [
    {
      "type": "tool_result",
      "tool_use_id": "toolu_01H7JQ",
      "content": [{ "type": "text", "text": "18 C and clear" }]
    },
    { "type": "text", "text": "Summarize that in one sentence." }
  ]
}
```

`tool_choice` accepts these forms:

| Value | Behavior |
| - | - |
| `{ "type": "auto" }` | The model decides whether to call a tool. |
| `{ "type": "none" }` | The model does not call a tool. |
| `{ "type": "tool", "name": "get_weather" }` | The model must call the named tool. |
| `{ "type": "any" }` | The model must call a tool. On models whose page says forced tool calling is limited, this value with several tools is refused. Use `auto` or name one tool. |

Add `"disable_parallel_tool_use": true` inside `tool_choice` when the model must request at most one tool in that turn.

Anthropic server tools, including bash, text editor, web search, and computer use, are refused. This route supports client tools that your application executes.

## Thinking and structured output

Set `thinking.type` to `enabled` or `adaptive` to keep model reasoning on. On DeepSeek V4 Pro, either value turns reasoning on. Set it to `disabled` to turn reasoning off on a model that supports disabling it.

Some models cannot disable reasoning. The request still succeeds, the setting is ignored, and `X-RunInfra-Hint-Warnings` says `thinking cannot be disabled on this model`. The documented behavior for GLM 5.3 Flash and Qwen3.8 2.4T A95B follows this rule. Check current availability before selecting either model.

`thinking.budget_tokens` is accepted and ignored. Reasoning spends the same `max_tokens` budget as the final answer.

Use `output_config.effort` with `none`, `minimal`, `low`, `medium`, `high`, `xhigh`, or `max` to select the reasoning effort. A value the model does not support can be moved to the nearest one it does, and `X-RunInfra-Hint-Warnings` then carries `reasoning_effort_clamped`. Use `output_config.format` with `type: "json_schema"` to request structured output.

## Accepted compatibility fields

These Anthropic-only fields are accepted and ignored: `cache_control`, `service_tier`, `container`, `mcp_servers`, `context_management`, `inference_geo`, and `speed`. They do not change routing, billing, or model behavior.

## Response

The `content` array preserves this order when the parts exist: a `thinking` block, a `text` block, then `tool_use` blocks.

```json theme={"dark"}
{
  "id": "msg_a1b2c3d4",
  "type": "message",
  "role": "assistant",
  "content": [
    {
      "type": "thinking",
      "thinking": "I should check the weather tool.",
      "signature": "opaque_runinfra_signature"
    },
    {
      "type": "text",
      "text": "I will check the current weather."
    },
    {
      "type": "tool_use",
      "id": "toolu_01H7JQ",
      "name": "get_weather",
      "input": { "city": "Amman" }
    }
  ],
  "model": "nemotron-3-5-lightning-30b",
  "stop_reason": "tool_use",
  "stop_sequence": null,
  "usage": {
    "input_tokens": 24,
    "output_tokens": 31,
    "cache_read_input_tokens": 128,
    "cache_creation_input_tokens": 0
  }
}
```

The thinking `signature` is opaque. It is not an Anthropic attestation. You may send it back in a prior assistant turn, where it is accepted and dropped.

| Field | Behavior |
| - | - |
| `stop_reason` | `end_turn`, `max_tokens`, `tool_use`, or `null`. |
| `stop_sequence` | Always `null`. The model honors a matching sequence, but the response does not report which one matched. |
| `usage.input_tokens` | Prompt tokens not served from cache. |
| `usage.cache_read_input_tokens` | Prompt-prefix tokens served from cache. |
| `usage.cache_creation_input_tokens` | Always `0`. |
| `usage.output_tokens` | All generated output tokens, including reasoning. |

Your billed prompt count is `input_tokens + cache_read_input_tokens`.

## Count input tokens

```http theme={"dark"}
POST https://api.runinfra.ai/v1/messages/count_tokens
```

Use the same body as `POST /v1/messages`. `max_tokens` is optional for this operation. The endpoint estimates input tokens without running the model, so the call is not billed. The per-key request rate limit still applies.

<CodeGroup>
  ```python Python theme={"dark"}
  count = client.messages.count_tokens(
      model="nemotron-3-5-lightning-30b",
      messages=[{"role": "user", "content": "Summarize this design."}],
  )

  print(count.input_tokens)
  ```

  ```typescript TypeScript theme={"dark"}
  const count = await client.messages.countTokens({
    model: "nemotron-3-5-lightning-30b",
    messages: [{ role: "user", content: "Summarize this design." }],
  });

  console.log(count.input_tokens);
  ```

  ```bash cURL theme={"dark"}
  curl https://api.runinfra.ai/v1/messages/count_tokens \
    -H "x-api-key: $RUNINFRA_GATEWAY_KEY" \
    -H "anthropic-version: 2023-06-01" \
    -H "Content-Type: application/json" \
    -d '{"model":"nemotron-3-5-lightning-30b","messages":[{"role":"user","content":"Summarize this design."}]}'
  ```
</CodeGroup>

```json theme={"dark"}
{
  "input_tokens": 12
}
```

| Field | Behavior |
| - | - |
| `input_tokens` | Estimated input tokens for the translated request. |

## Streaming

Set `stream: true` to receive event-named Server-Sent Events. There is no `data: [DONE]` sentinel.

Blocks open only when their first output arrives:

* A thinking block emits `thinking_delta` fragments, then one `signature_delta`.
* A text block emits `text_delta` fragments.
* A tool-use block opens with `input: {}`, then emits `input_json_delta` fragments. Concatenate each `partial_json` value to recover the full input object.

The stream order is `message_start`, lazily opened content blocks, `message_delta`, then `message_stop`. The server emits `ping` every 15 seconds while the model is silent.

```text theme={"dark"}
event: message_start
data: {"type":"message_start","message":{"id":"msg_...","type":"message","role":"assistant","model":"nemotron-3-5-lightning-30b","content":[],"stop_reason":null,"stop_sequence":null,"usage":{"input_tokens":0,"output_tokens":0,"cache_read_input_tokens":0,"cache_creation_input_tokens":0}}}

event: content_block_start
data: {"type":"content_block_start","index":0,"content_block":{"type":"thinking","thinking":"","signature":""}}

event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"thinking_delta","thinking":"I should call "}}

event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"signature_delta","signature":"opaque_runinfra_signature"}}

event: content_block_stop
data: {"type":"content_block_stop","index":0}

event: content_block_start
data: {"type":"content_block_start","index":1,"content_block":{"type":"tool_use","id":"toolu_01H7JQ","name":"get_weather","input":{}}}

event: content_block_delta
data: {"type":"content_block_delta","index":1,"delta":{"type":"input_json_delta","partial_json":"{\"city\":\"Amman\"}"}}

event: content_block_stop
data: {"type":"content_block_stop","index":1}

event: message_delta
data: {"type":"message_delta","delta":{"stop_reason":"tool_use","stop_sequence":null},"usage":{"input_tokens":24,"output_tokens":31,"cache_read_input_tokens":128,"cache_creation_input_tokens":0}}

event: message_stop
data: {"type":"message_stop"}
```

`message_start` reports zero usage. The cumulative four-field usage on `message_delta` is authoritative.

The official SDK stream helpers work unchanged, including Python `messages.stream()` and `get_final_message()`, and TypeScript `messages.stream()` and `finalMessage()`.

If a failure happens after the stream opens, the final frame is an `error` event and there is no `message_stop`:

```text theme={"dark"}
event: error
data: {"type":"error","error":{"type":"api_error","message":"The model connection closed before completion."},"request_id":"req_..."}
```

## Errors

Every `4xx` and `5xx` on the two Anthropic-compatible routes uses this envelope:

```json theme={"dark"}
{
  "type": "error",
  "error": {
    "type": "invalid_request_error",
    "message": "max_tokens must be at least 1."
  },
  "request_id": "req_..."
}
```

Messages match the corresponding chat-completions errors. Branch on `error.type`, not the message text. A prompt longer than the model's context window is the exception: its `invalid_request_error` message starts `prompt is too long`, the wording Claude Code compacts on.

| Status | `error.type` | Typical cause |
| -: | - | - |
| `400` | `invalid_request_error` | Invalid JSON, a missing field, an unsupported content block, or an unsupported server tool. |
| `401` | `authentication_error` | The key is missing or invalid. |
| `402` | `billing_error` | A billing refusal, including insufficient credits, an account hold, or a coding plan limit. Read the message for the reason and next step. |
| `403` | `permission_error` | The key cannot use the selected model. |
| `404` | `not_found_error` | The model is not available to this key. |
| `409` | `conflict_error` | The same idempotency key is still running. |
| `413` | `request_too_large` | The encoded body exceeds the request-size limit. |
| `429` | `rate_limit_error` | A request, token, or concurrency limit refused the call. Read `Retry-After`. |
| `500`, `502` | `api_error` | A transient gateway or model failure. |
| `503` | `overloaded_error` | The service or selected model is temporarily unavailable. The same type is returned when an `Idempotency-Key` could not be checked: the message says so, the request did not run, and nothing was charged. Retry with the same key after `Retry-After`. |
| `504` | `timeout_error` | The inference deadline expired. |

## Coding plan funding

A covered Messages request follows the same plan, credits, and Standby order as chat completions. [Funding headers](/docs/api-reference/usage#response-funding-headers) identify the request's admission decision. When the plan or Standby pays, Messages success usage also carries `usage.cost` (always `0`) and `usage.runinfra` with `cost_microcents`, `plan_value_microcents` and `paid_from`, in the JSON body and in the final `message_delta` of a stream. Credit-funded and pay-as-you-go Messages usage carries the four token counters only, so read `x-runinfra-funding` to see that credits paid after a limit.

A plan limit returns HTTP `402`, `x-should-retry: false`, and no `Retry-After`. The error type is `billing_error`. Only the message text survives from the detailed plan refusal, so it carries the product, limit, reset time, Standby condition, and next step with its Billing URL. The envelope does not carry `error.code`, `standby_blocker`, credit counters, or `fixes`.

In Messages clients such as Claude Code, read the message for the reason and next step. The CLI's connection check preserves that self-contained refusal text even when the envelope has no error code. Display the message and read the funding headers. Use [GET /v1/usage](/docs/api-reference/usage) for structured current plan state. Do not try to read the OpenAI-compatible [plan limit fields](/docs/api-reference/errors#coding-plan-limit) from a Messages error.

## Related

<Columns cols={3}>
  <Card title="Claude Code" icon="terminal" href="/docs/tools-sdks/claude-code">
    Configure Claude Code against the Anthropic-compatible routes.
  </Card>

  <Card title="Authentication" icon="key" href="/docs/api-reference/authentication">
    Choose the correct header and base URL.
  </Card>

  <Card title="OpenAI compatibility" icon="plug" href="/docs/tools-sdks/openai-compatibility">
    Use chat completions or Responses with OpenAI clients.
  </Card>
</Columns>
