Skip to main content
This is the endpoint every Model APIs caller reaches. Send a standard chat completion from the SDK you already use, with one of the live model ids from GET /v1/models.

What each model supports

The model-specific behavior and dated measurements below also cover models that may be paused. A capability does not imply current availability.
Run GET /v1/models for the current model list and each model’s capabilities. A model that is temporarily paused is reported there and on its Model Library page.
Requesting a capability a model does not have returns 400 with error.code hosted_capability_not_supported and param naming the field. You find out at the call, not in the output.

Give reasoning models room

Reasoning models can spend their output budget before they answer. Nemotron 3.5 Lightning 30B reasons by default. Reasoning tokens count toward max_tokens before the answer, so set max_tokens to at least the figure in the table below. A smaller budget can be spent entirely on reasoning and return a completion whose content is empty.
Reasoning models: output budgetreasoning tokens count toward max_tokensmax_tokens: 2048budget exhaustedThe whole budget went to reasoning: no answer. No-answer responses now settle at zero.max_tokens: 16384answer deliveredreasoninganswerEach reasoning model’s page publishes its recommended minimum.Measured 2026-08-14: at 2,048 tokens both public models returned an empty or truncated answer.Those probes billed output. DeepSeek V4 Flash: use 16,384. Segment widths are illustrative.
Reasoning models: output budgetreasoning tokens count toward max_tokensmax_tokens: 2048budget exhaustedThe whole budget went to reasoning: no answer. No-answer responses now settle at zero.max_tokens: 16384answer deliveredreasoninganswerEach reasoning model’s page publishes its recommended minimum.Measured 2026-08-14: at 2,048 tokens both public models returned an empty or truncated answer.Those probes billed output. DeepSeek V4 Flash: use 16,384. Segment widths are illustrative.
Measured on this endpoint on 2026-08-14 by sending the prompt “Write a TypeScript retry helper.” and recording how often a complete answer came back: For another model, read the recommended output budget on its page in the Model Library. The budget is a ceiling, not a spend. At 16,384 DeepSeek V4 Flash still stopped on its own after roughly 7,500 output tokens, and raising the ceiling further did not make it generate more. A ceiling below the table’s figure produces the reasoning and not the answer.
RunInfra preserves finish_reason: "length" and adds choices[].runinfra.output_status when a choice reaches its generation limit before producing final answer content.
Branch on choices[].runinfra.output_status.code. generation_limit_reached_before_answer means the choice ended for length. If another terminal finish_reason produced no answer content, the code is no_answer_content and you should inspect finish_reason before retrying.The aggregate usage annotation appears only when every choice lacks a final answer. RunInfra does not estimate a reasoning-token count: completion_tokens stays the model-reported output total, and non_answer_completion_tokens records that those tokens produced no answer. The annotation changes no pricing; a response with no billable output settles at zero, as described below.

The cost of the request, and its cached input, on its usage

Every hosted model response carries the calculated charge for the request. On credits, it uses the settlement formula over the tokens the provider reported. On a coding plan or Standby, the calculated charge is zero and the usage value is reported separately. A response with no billable output settles at zero and prints 0. Beside the cost, every hosted response carries the count of input tokens that were billed at the cached input rate.
The cost covers the tokens the provider reported, at the model’s prices for uncached input, cached input (when the model discloses a cache hit) and output. During a promotional free window it is 0, present, not absent. These cost and cache fields describe hosted models. The cached count is the number the cost was computed with. On a billable response, prompt_tokens - cached_tokens is exactly what was billed at the uncached input rate. The cached count is 0 rather than absent whenever nothing was billed as cached: a first send, a prompt too short to match, a response that settled at zero, or a model whose cache is shared across tenants. On a shared cache a hit is neither disclosed nor priced, so 0 there means no cached discount was applied, not that the cache did nothing. /v1/models says which models publish an isolated cache. cost_microcents is the canonical figure and cost is the same amount expressed in dollars. Your balance is debited in whole cents, and the fraction of a cent left over carries to your next request, so one debit can differ from one request’s cost by less than a cent. Over any period the two agree to within one cent of carry. On a stream, the cost and the cached count ride on the usage frame that is sent together with [DONE], so send stream_options: {"include_usage": true} to receive them (the one usage frame forwarded without it, when a stream produced no answer, carries them too). Content chunks carry neither; a running figure on every delta would be a guess. A stream that ends in an error never quotes a cost. A /v1/responses usage carries the cost in the same two fields, and the cached count as input_tokens_details.cached_tokens and runinfra.cached_input_tokens. The usage of an idempotent replay of a streamed request (X-RunInfra-Idempotent-Replay: true) carries the figures that were settled. Read the funding headers for the request’s admission decision. Plan usage reports current plan limits and policy; Credits and budget reports the credit balance.

Turning reasoning off

reasoning_effort: "none" turns reasoning off for one request. The completion then carries the answer and no reasoning stream. Measured on this endpoint on 2026-08-16 with an identical prompt at temperature 0, the completion fell from 35 tokens to 3 on DeepSeek V4 Flash, from 335 to 5 on nemotron-3-5-lightning-30b, and from 67 to 5 on qwen3-8-27b. On Qwen3.8 Flash Next, measured 2026-08-29 on a one-line prompt, the completion fell from 30 tokens at the default effort to 2 with none. Because reasoning bills at the output rate, this is the largest per-request cost lever on short tasks. Two models cannot turn reasoning off. Qwen3.8 2.4T A95B and GLM 5.3 Flash both refuse reasoning_effort: "none" with 400 rather than accepting it and reasoning anyway, so the request is never billed for effort you asked to skip. Both refusals name the smallest budget they accept. Send "low" for the smallest reasoning budget on either.
Measured on 2026-08-16 at temperature 0 with repeated runs per level, with the later measurements dated below.
  • DeepSeek V4 Flash applies max when you send no reasoning_effort, and treats high and xhigh as max. An omitted field on a request that can call tools (tools sent and tool_choice not "none") without a JSON response_format runs at medium. medium is reproducibly distinct from the default.
  • DeepSeek V4 Pro answers without reasoning at the default, so there is nothing to turn off; the levels are not measured.
  • Qwen3.8 2.4T A95B produces reproducibly distinct reasoning at low, medium and xhigh (113, 140 and 89 completion tokens on the probe, with xhigh equal to the omitted default). high is sent as xhigh. none is refused.
  • qwen3-8-27b measurably alters its reasoning at medium, against a twice-identical baseline. high and max are sent as xhigh, the model’s maximum and its default.
  • nemotron-3-5-lightning-30b level deltas stayed inside the model’s own run-to-run variance, so treat effort there as the off switch only.
  • GLM 5.3 Flash: low and high are the two tiers the model honors as sent. An omitted field and every other accepted value, medium included, run at the model’s maximum effort. none is refused. Send low for the smallest budget. Measured on 2026-08-27.
  • ornith-1-5-35b reasons at the default and none turns it off; the levels between are not measured.
  • Qwen3.8 Flash Next: none is accepted and returns the answer directly. high and max are sent as xhigh, the model’s maximum and its default. low and medium are accepted as sent. minimal is refused with 400 and param: "reasoning_effort". Measured on 2026-08-30.
The vendor spellings enable_thinking, chat_template_kwargs.enable_thinking, thinking_budget and thinking are accepted for request-shape compatibility and have no effect on any listed model: a request relying on them runs at the model’s default effort and can still produce billable reasoning. Send reasoning_effort instead; it is the one honored spelling.

What the gateway does with a field you send

What happens to a field you sendPOST /v1/chat/completionsChecked, then sentThe gateway validates the typeor the range first. A bad valueis refused with 400 before therequest is served or billed.model, messages, max_tokensSent untouchedOn the forwarding allowlist.Any JSON value, no gatewaycheck. The model can stillreject it, as a 400.seed, user, service_tierRemovedNot on the allowlist, so it isdropped before forwarding. Themodel never sees it and youget no warning.top_k, guided_json, reasoningThe allowlist is exhaustive. An unknown top-level field passes request validation and is still removedbefore the request reaches the model. Check the list before depending on a field.
What happens to a field you sendPOST /v1/chat/completionsChecked, then sentThe gateway validates the typeor the range first. A bad valueis refused with 400 before therequest is served or billed.model, messages, max_tokensSent untouchedOn the forwarding allowlist.Any JSON value, no gatewaycheck. The model can stillreject it, as a 400.seed, user, service_tierRemovedNot on the allowlist, so it isdropped before forwarding. Themodel never sees it and youget no warning.top_k, guided_json, reasoningThe allowlist is exhaustive. An unknown top-level field passes request validation and is still removedbefore the request reaches the model. Check the list before depending on a field.
Every field lands in one of three dispositions. These are checked against the contract below, then forwarded: A message’s role is system, developer, user, assistant or tool. Any other role returns 400. content may be a string, an array, or null. An assistant message needs content or a non-empty tool_calls. Each tool_calls[].function.arguments must hold a JSON object, such as "{}" for a call with no arguments. A tool message needs content and a non-empty string tool_call_id. Every other role needs content.

System message placement

A developer message is sent to the model as a system message. Most models accept system messages anywhere in the conversation and receive them where you put them. Qwen3.8 27B, Qwen3.8 2.4T A95B, Qwen3.8 Flash Next, and Ornith 1.5 35B accept one system message, and only as the first message. For these models the gateway repairs the order before the request reaches the model:
  • Consecutive system messages at the start are merged into one, in order, separated by a blank line.
  • A system message later in the conversation is sent as a user message in the same position.
Every other message, tool call, and tool result keeps its position and content. A repaired conversation still extends the previous turn’s prompt, so prompt caching keeps working across turns. Coding agents that send more than one system message, such as Claude Code and Codex CLI, run on these models without extra configuration. A leading system message that contains anything other than text is not merged, and the request returns 400 with param: "messages".
The gateway accepts any JSON value on these, applies no default, and forwards it unchanged. The model can still reject the value, which comes back as 400 invalid_request_error.audio, frequency_penalty, function_call, functions, logit_bias, logprobs, metadata, modalities, moderation, parallel_tool_calls, prediction, presence_penalty, safety_identifier, seed, service_tier, store, tool_choice, top_logprobs, user, verbosity, web_search_options.Some allowlisted fields do carry gateway behavior:
The allowlist is exhaustive: every other top-level field is dropped. Common examples are max_output_tokens, reasoning, text, include, truncation and previous_response_id (Responses API fields on the Chat route); top_k, min_p, repetition_penalty and length_penalty; any nonstandard guided-decoding or beam-search field such as guided_json, guided_regex, guided_choice, guided_grammar and use_beam_search (use response_format and tools for constrained output); and num_return_sequences, num_beams and beam_width, which the gateway reads for the four-choice safety check and then removes.An unknown field can pass request validation and still be removed here, so never rely on silent pass-through for a field that is not on the allowlist.

Output limits

The body-size and image limits count every message in the request, including the images and history a client resends each turn. A single image fits at about 2.6MB on disk because base64 adds a third; drop older images from the conversation or split them across requests. Omit both output-token fields and the gateway forwards no output limit at all, so the model generates until it stops on its own or reaches the remaining context. Send an oversized max_tokens and the gateway clamps it to what the context allows rather than refusing the request: the response carries X-RunInfra-Output-Clamped: true and X-RunInfra-Output-Token-Ceiling with the value used. Two bounds apply at once and the smaller one wins. Tokens are bounded by the remaining context above; wall-clock is bounded by the 740 second response limit, which at typical decode rates is the binding limit for very long generations. /v1/models publishes both, max_output_tokens beside response_time_ceiling_seconds, which carries the same 740 second figure, so size a long generation against the pair rather than the token figure alone. The 504 message names the budget that was exceeded. For a non-streaming request, the error message suggests two remedies: stream: true (response headers arrive with the first token) and a smaller max_tokens. A first-token timeout on a stream suggests a smaller prompt, or retrying when the model has more capacity. A model’s context window is a separate per-model limit, published on that model’s page in the Model Library. Served context is a property of how the model is served, not of the weights, so read it on the model page rather than a model card.

Three things a request can be refused for

Each one is refused up front with a 400 naming the field, rather than answered with output you cannot use.
response_format of json_schema or json_object is enforced during generation, and on some models that enforcement cannot start until the model has finished reasoning. On those models, omitting reasoning_effort applies an effort measured to return a conforming object: reasoning_effort: "none" on DeepSeek V4 Flash and nemotron-3-5-lightning-30b, and "low" on GLM 5.3 Flash, where an explicit "high" also conforms. The gateway reports an automatically applied value in the X-RunInfra-Reasoning-Effort-Applied response header. Send an incompatible effort beside the format and the request is refused with 400, code hosted_parameter_not_supported and param: "response_format".The fix is the effort value sent alongside response_format:
Check each model’s JSON mode requirements in the Model Library before changing the model in this request. response_format: {"type": "text"} constrains nothing and is never refused. On /v1/responses the same remedy is a top-level reasoning_effort beside text.format. Working snippets are in Structured output.
Accepted input is a per-model fact, stated on each model page as the Accepted input row. Qwen3.8 27B, Ornith 1.5 35B, GLM 5.3 Flash and Qwen3.8 Flash Next accept image_url and input_image content parts as inline data URLs, up to 8 images per request: {"type": "image_url", "image_url": {"url": "data:image/png;base64,..."}} with png, jpeg, webp or gif. Remote image URLs are not fetched; download the image, base64-encode it, and send the data URL. Image tokens are counted by the server inside prompt_tokens and bill at the model’s input rate. Each image must be a complete file: a png cut inside one of its data chunks, a jpeg without its end-of-image marker, or a webp shorter than its header declares is refused with a 400 that names the message and part (for example messages[2].content[1]) before any generation starts, and an image that still cannot be decoded is refused the same way rather than reported as unavailable.On every other live model, and for audio and video everywhere, a non-text content part inside messages returns 400 with param: "messages" and code: "hosted_parameter_not_supported", naming the model and the part to remove. The part is refused rather than stripped, so a request carrying media is never billed as if it were text.
prompt_cache_options and prompt_cache_retention are rejected with 400 hosted_parameter_not_supported. prompt_cache_key is accepted: when no session header supplies an identity, the gateway reads it as a cache-affinity hint and removes it before the request reaches the model. It grants no control over retention. Prefix caching, where a model has it, is automatic server behavior and is not controllable per request.

Prefix cache retention

Retention is what GET /v1/models states as cache_tiers and cache_retention. DeepSeek V4 Pro reports cache_retention: "tiered_host_memory": an idle session’s cached prefix survives beyond immediate memory pressure and can be reused when the session resumes. Every other chat model, GLM 5.3 Flash and Qwen3.8 Flash Next included, reports "best_effort": cached prefixes are retained opportunistically and may be evicted at any time, with no retention promise. Keeping a session’s follow-up turns close to its warm cache is covered in Session affinity; read usage.prompt_tokens_details.cached_tokens for what the cache did on each request.

Calling a paused model

A hosted model can be temporarily paused. The request returns 503 with code hosted_model_paused and a Retry-After header carrying the seconds until availability is next rechecked, capped at 3,600, or 60 when no recheck time is known or the known one has already passed. error.paused_until names that next recheck time, present only when one is scheduled; it is a recheck deadline, not a promised return. Nothing is charged. The model id stays valid and stays listed by GET /v1/models with availability: "paused", so retry rather than re-resolving your configuration.

Streaming

Read deltas and ask for the usage frame.

Rate limits

The four layers that can refuse you, and how to back off.

Errors

Every status, code, and caller action.