GET /v1/models.
What each model supports
The model-specific behavior and dated measurements below also cover models that may be paused. A capability does not imply current availability.Run
GET /v1/models for the current model list and each model’s capabilities. A model that is temporarily paused is reported there and on its Model Library page.400 with error.code hosted_capability_not_supported and param naming the field. You find out at the call, not in the output.
Give reasoning models room
Reasoning models can spend their output budget before they answer. Nemotron 3.5 Lightning 30B reasons by default. Reasoning tokens count towardmax_tokens before the answer, so set max_tokens to at least the figure in the table below. A smaller budget can be spent entirely on reasoning and return a completion whose content is empty.
Measured on this endpoint on 2026-08-14 by sending the prompt “Write a TypeScript retry helper.” and recording how often a complete answer came back:
For another model, read the recommended output budget on its page in the Model Library.
The budget is a ceiling, not a spend. At 16,384 DeepSeek V4 Flash still stopped on its own after roughly 7,500 output tokens, and raising the ceiling further did not make it generate more. A ceiling below the table’s figure produces the reasoning and not the answer.
How to detect an exhausted budget in code
How to detect an exhausted budget in code
RunInfra preserves Branch on
finish_reason: "length" and adds choices[].runinfra.output_status when a choice reaches its generation limit before producing final answer content.choices[].runinfra.output_status.code. generation_limit_reached_before_answer means the choice ended for length. If another terminal finish_reason produced no answer content, the code is no_answer_content and you should inspect finish_reason before retrying.The aggregate usage annotation appears only when every choice lacks a final answer. RunInfra does not estimate a reasoning-token count: completion_tokens stays the model-reported output total, and non_answer_completion_tokens records that those tokens produced no answer. The annotation changes no pricing; a response with no billable output settles at zero, as described below.The cost of the request, and its cached input, on its usage
Every hosted model response carries the calculated charge for the request. On credits, it uses the settlement formula over the tokens the provider reported. On a coding plan or Standby, the calculated charge is zero and the usage value is reported separately. A response with no billable output settles at zero and prints0. Beside the cost, every hosted response carries the count of input tokens that were billed at the cached input rate.
The cost covers the tokens the provider reported, at the model’s prices for uncached input, cached input (when the model discloses a cache hit) and output. During a promotional free window it is
0, present, not absent. These cost and cache fields describe hosted models.
The cached count is the number the cost was computed with. On a billable response, prompt_tokens - cached_tokens is exactly what was billed at the uncached input rate. The cached count is 0 rather than absent whenever nothing was billed as cached: a first send, a prompt too short to match, a response that settled at zero, or a model whose cache is shared across tenants. On a shared cache a hit is neither disclosed nor priced, so 0 there means no cached discount was applied, not that the cache did nothing. /v1/models says which models publish an isolated cache.
cost_microcents is the canonical figure and cost is the same amount expressed in dollars. Your balance is debited in whole cents, and the fraction of a cent left over carries to your next request, so one debit can differ from one request’s cost by less than a cent. Over any period the two agree to within one cent of carry.
On a stream, the cost and the cached count ride on the usage frame that is sent together with [DONE], so send stream_options: {"include_usage": true} to receive them (the one usage frame forwarded without it, when a stream produced no answer, carries them too). Content chunks carry neither; a running figure on every delta would be a guess. A stream that ends in an error never quotes a cost. A /v1/responses usage carries the cost in the same two fields, and the cached count as input_tokens_details.cached_tokens and runinfra.cached_input_tokens. The usage of an idempotent replay of a streamed request (X-RunInfra-Idempotent-Replay: true) carries the figures that were settled.
Read the funding headers for the request’s admission decision. Plan usage reports current plan limits and policy; Credits and budget reports the credit balance.
Turning reasoning off
reasoning_effort: "none" turns reasoning off for one request. The completion then carries the answer and no reasoning stream. Measured on this endpoint on 2026-08-16 with an identical prompt at temperature 0, the completion fell from 35 tokens to 3 on DeepSeek V4 Flash, from 335 to 5 on nemotron-3-5-lightning-30b, and from 67 to 5 on qwen3-8-27b. On Qwen3.8 Flash Next, measured 2026-08-29 on a one-line prompt, the completion fell from 30 tokens at the default effort to 2 with none. Because reasoning bills at the output rate, this is the largest per-request cost lever on short tasks.
Two models cannot turn reasoning off. Qwen3.8 2.4T A95B and GLM 5.3 Flash both refuse reasoning_effort: "none" with 400 rather than accepting it and reasoning anyway, so the request is never billed for effort you asked to skip. Both refusals name the smallest budget they accept. Send "low" for the smallest reasoning budget on either.
How the effort levels behave between none and the ceiling
How the effort levels behave between none and the ceiling
Measured on 2026-08-16 at temperature 0 with repeated runs per level, with the later measurements dated below.
- DeepSeek V4 Flash applies
maxwhen you send noreasoning_effort, and treatshighandxhighasmax. An omitted field on a request that can call tools (toolssent andtool_choicenot"none") without a JSONresponse_formatruns atmedium.mediumis reproducibly distinct from the default. - DeepSeek V4 Pro answers without reasoning at the default, so there is nothing to turn off; the levels are not measured.
- Qwen3.8 2.4T A95B produces reproducibly distinct reasoning at
low,mediumandxhigh(113, 140 and 89 completion tokens on the probe, withxhighequal to the omitted default).highis sent asxhigh.noneis refused. qwen3-8-27bmeasurably alters its reasoning atmedium, against a twice-identical baseline.highandmaxare sent asxhigh, the model’s maximum and its default.nemotron-3-5-lightning-30blevel deltas stayed inside the model’s own run-to-run variance, so treat effort there as the off switch only.- GLM 5.3 Flash:
lowandhighare the two tiers the model honors as sent. An omitted field and every other accepted value,mediumincluded, run at the model’s maximum effort.noneis refused. Sendlowfor the smallest budget. Measured on 2026-08-27. ornith-1-5-35breasons at the default andnoneturns it off; the levels between are not measured.- Qwen3.8 Flash Next:
noneis accepted and returns the answer directly.highandmaxare sent asxhigh, the model’s maximum and its default.lowandmediumare accepted as sent.minimalis refused with400andparam: "reasoning_effort". Measured on 2026-08-30.
enable_thinking, chat_template_kwargs.enable_thinking, thinking_budget and thinking are accepted for request-shape compatibility and have no effect on any listed model: a request relying on them runs at the model’s default effort and can still produce billable reasoning. Send reasoning_effort instead; it is the one honored spelling.What the gateway does with a field you send
Every field lands in one of three dispositions. These are checked against the contract below, then forwarded:
A message’s
role is system, developer, user, assistant or tool. Any other role returns 400. content may be a string, an array, or null. An assistant message needs content or a non-empty tool_calls. Each tool_calls[].function.arguments must hold a JSON object, such as "{}" for a call with no arguments. A tool message needs content and a non-empty string tool_call_id. Every other role needs content.
System message placement
Adeveloper message is sent to the model as a system message. Most models accept system messages anywhere in the conversation and receive them where you put them.
Qwen3.8 27B, Qwen3.8 2.4T A95B, Qwen3.8 Flash Next, and Ornith 1.5 35B accept one system message, and only as the first message. For these models the gateway repairs the order before the request reaches the model:
- Consecutive system messages at the start are merged into one, in order, separated by a blank line.
- A system message later in the conversation is sent as a
usermessage in the same position.
400 with param: "messages".
Sent untouched: the forwarding allowlist
Sent untouched: the forwarding allowlist
The gateway accepts any JSON value on these, applies no default, and forwards it unchanged. The model can still reject the value, which comes back as
400 invalid_request_error.audio, frequency_penalty, function_call, functions, logit_bias, logprobs, metadata, modalities, moderation, parallel_tool_calls, prediction, presence_penalty, safety_identifier, seed, service_tier, store, tool_choice, top_logprobs, user, verbosity, web_search_options.Some allowlisted fields do carry gateway behavior:Removed before the request leaves the gateway
Removed before the request leaves the gateway
The allowlist is exhaustive: every other top-level field is dropped. Common examples are
max_output_tokens, reasoning, text, include, truncation and previous_response_id (Responses API fields on the Chat route); top_k, min_p, repetition_penalty and length_penalty; any nonstandard guided-decoding or beam-search field such as guided_json, guided_regex, guided_choice, guided_grammar and use_beam_search (use response_format and tools for constrained output); and num_return_sequences, num_beams and beam_width, which the gateway reads for the four-choice safety check and then removes.An unknown field can pass request validation and still be removed here, so never rely on silent pass-through for a field that is not on the allowlist.Output limits
The body-size and image limits count every message in the request, including the images and history a client resends each turn. A single image fits at about 2.6MB on disk because base64 adds a third; drop older images from the conversation or split them across requests.
Omit both output-token fields and the gateway forwards no output limit at all, so the model generates until it stops on its own or reaches the remaining context. Send an oversized
max_tokens and the gateway clamps it to what the context allows rather than refusing the request: the response carries X-RunInfra-Output-Clamped: true and X-RunInfra-Output-Token-Ceiling with the value used.
Two bounds apply at once and the smaller one wins. Tokens are bounded by the remaining context above; wall-clock is bounded by the 740 second response limit, which at typical decode rates is the binding limit for very long generations. /v1/models publishes both, max_output_tokens beside response_time_ceiling_seconds, which carries the same 740 second figure, so size a long generation against the pair rather than the token figure alone.
The 504 message names the budget that was exceeded. For a non-streaming request, the error message suggests two remedies: stream: true (response headers arrive with the first token) and a smaller max_tokens. A first-token timeout on a stream suggests a smaller prompt, or retrying when the model has more capacity.
A model’s context window is a separate per-model limit, published on that model’s page in the Model Library. Served context is a property of how the model is served, not of the weights, so read it on the model page rather than a model card.
Three things a request can be refused for
Each one is refused up front with a400 naming the field, rather than answered with output you cannot use.
A JSON format the model cannot hold while reasoning
A JSON format the model cannot hold while reasoning
response_format of json_schema or json_object is enforced during generation, and on some models that enforcement cannot start until the model has finished reasoning. On those models, omitting reasoning_effort applies an effort measured to return a conforming object: reasoning_effort: "none" on DeepSeek V4 Flash and nemotron-3-5-lightning-30b, and "low" on GLM 5.3 Flash, where an explicit "high" also conforms. The gateway reports an automatically applied value in the X-RunInfra-Reasoning-Effort-Applied response header. Send an incompatible effort beside the format and the request is refused with 400, code hosted_parameter_not_supported and param: "response_format".The fix is the effort value sent alongside response_format:response_format: {"type": "text"} constrains nothing and is never refused. On /v1/responses the same remedy is a top-level reasoning_effort beside text.format. Working snippets are in Structured output.An image, audio, or video content part
An image, audio, or video content part
Accepted input is a per-model fact, stated on each model page as the Accepted input row. Qwen3.8 27B, Ornith 1.5 35B, GLM 5.3 Flash and Qwen3.8 Flash Next accept
image_url and input_image content parts as inline data URLs, up to 8 images per request: {"type": "image_url", "image_url": {"url": "data:image/png;base64,..."}} with png, jpeg, webp or gif. Remote image URLs are not fetched; download the image, base64-encode it, and send the data URL. Image tokens are counted by the server inside prompt_tokens and bill at the model’s input rate. Each image must be a complete file: a png cut inside one of its data chunks, a jpeg without its end-of-image marker, or a webp shorter than its header declares is refused with a 400 that names the message and part (for example messages[2].content[1]) before any generation starts, and an image that still cannot be decoded is refused the same way rather than reported as unavailable.On every other live model, and for audio and video everywhere, a non-text content part inside messages returns 400 with param: "messages" and code: "hosted_parameter_not_supported", naming the model and the part to remove. The part is refused rather than stripped, so a request carrying media is never billed as if it were text.A prompt-cache control field
A prompt-cache control field
prompt_cache_options and prompt_cache_retention are rejected with 400 hosted_parameter_not_supported. prompt_cache_key is accepted: when no session header supplies an identity, the gateway reads it as a cache-affinity hint and removes it before the request reaches the model. It grants no control over retention. Prefix caching, where a model has it, is automatic server behavior and is not controllable per request.Prefix cache retention
Retention is whatGET /v1/models states as cache_tiers and cache_retention. DeepSeek V4 Pro reports cache_retention: "tiered_host_memory": an idle session’s cached prefix survives beyond immediate memory pressure and can be reused when the session resumes. Every other chat model, GLM 5.3 Flash and Qwen3.8 Flash Next included, reports "best_effort": cached prefixes are retained opportunistically and may be evicted at any time, with no retention promise. Keeping a session’s follow-up turns close to its warm cache is covered in Session affinity; read usage.prompt_tokens_details.cached_tokens for what the cache did on each request.
Calling a paused model
A hosted model can be temporarily paused. The request returns503 with code hosted_model_paused and a Retry-After header carrying the seconds until availability is next rechecked, capped at 3,600, or 60 when no recheck time is known or the known one has already passed. error.paused_until names that next recheck time, present only when one is scheduled; it is a recheck deadline, not a promised return. Nothing is charged. The model id stays valid and stays listed by GET /v1/models with availability: "paused", so retry rather than re-resolving your configuration.
Related
Streaming
Read deltas and ask for the usage frame.
Rate limits
The four layers that can refuse you, and how to back off.
Errors
Every status, code, and caller action.