Skip to main content
Model APIs and CLI updates, newest first. Older entries record what was available at the time. Prices, credits, and grants as they stand today are on Pricing and credits.
CLI
CLI 0.5.1: Goose, Zed and Doctor fixes
Run runinfra update to get the fixes.
  • On Windows, the command that loads the Goose or Zed key now works in a default PowerShell. If you set up Goose or Zed on Windows with 0.5.0, set it up again to get the new command.
  • Doctor no longer reports a missing key after a correct Goose, Zed, Mistral Vibe or Oh My Pi setup.
  • Shorter start steps for Goose and Zed.
  • The installer says that runinfra starts the step-by-step setup.
See CLI commands.
BillingCoding plan
A simpler Billing page
Settings > Billing now shows your balance, auto-recharge, your plan and invoices.
CLI
CLI 0.5.0: step-by-step terminal setup
Run runinfra in your terminal for step-by-step setup. Choose an agent and command, sign in in your browser, review the plan, and press Enter.
  • Use RunInfra inside Claude Code or Codex with --in-place; disconnect restores your earlier settings.
  • One review permits one apply attempt. Changed plans need a new review and your yes.
  • The tui, status, and launch commands are removed.
See CLI commands.
Coding planBilling
Coding plan: payment disputes and account deletion
After a payment dispute, a coding plan ends at its period end, even if the dispute is decided in RunInfra’s favor. Keep plan is no longer offered after a dispute, and can_keep stays false even after the hold is released. If the disputed card also bought credits, the account is frozen too, and buying credits does not lift that hold: contact support. A bank inquiry, which only asks about a charge, does not pause or end the plan. A renewal or an upgrade charged after you delete your account is refunded. See Coding plan.
Model APIs
Rate-limit store outages use per-server limits
A rate-limit store outage no longer refuses /v1 requests with 503 limiter_unavailable. Each API server applies the connection gate and per-key limits on its own, so X-RateLimit-* headers describe one server’s fixed 60 second window. Hosted admission and Idempotency-Key checks still need the store and can return a retryable 503. See Rate limits.
Coding planBillingAccount
Coding plan settings now live on the Coding plan page; Billing shows your plan at a glance; the Billing contact moved to Account

Coding plan settings on the Coding plan page

Settings > Coding plan is where an owner subscribes, changes or cancels a plan, right on the page. Subscribe, Change plan, Cancel plan and Change what happens at a limit open in place, with no dialog, and Not now closes them. Links to Billing’s coding plan section, from emails, the CLI or bookmarks, open this page.

Billing shows your plan at a glance

Settings > Billing shows your plan once, read-only, in its Plan and At a limit tiles. Its top band shows your balance alone. An owner can still pay an overdue amount or an unpaid upgrade there, and Manage in Coding plan and Cancel plan open the Coding plan page.

The Billing contact moved to Account

The Billing contact is now at Settings > Account, right after Account name. Only the owner sets it, and it works as before.See Manage your plan.
CLI
CLI 0.4.1: your coding agent can set up RunInfra for you
Your coding agent can set up RunInfra for you. Paste the prompt from Set up with your coding agent, approve sign-in, and agree to the setup plan.
  • Device sign-in returns a code and link without waiting. Your agent resumes the request after you select Approve. A pending approval has its own JSON result.
  • Connect returns the complete plan with exit 8; your agent adds --yes only after you agree. The result gives your agent the exact command, model picker instructions, and restart instructions to relay.
  • Supported agents add RunInfra next to your provider, use a separate command, or become the agent’s provider. Connections made with an earlier CLI keep their profiles. The agent guides explain each placement.
  • Choose Coding plan first or Pay as you go with --funding plan or --funding credits. The model list follows the key’s payer. Reconnecting can change that payer with owner approval; key rotation keeps it.

Terminal app

  • For people at a terminal, the app connects one agent per confirmation, shows the workspace and payer, gives the command, model picker, and restart instructions, and keeps temporary sessions, usage, and sign-out controls together. Coding agents use the JSON commands instead.
See CLI commands and Choose how to pay.
Security
Settings > Security: repeated rate-limit refusals folded into one event
Settings > Security folds repeated rate-limit refusals. When one key is refused many times within a minute, one event carries the count of the others. Its detail shows the count and when the refusals began, such as 499 more from 5m ago, on desktop and phone. When the start time is unknown, the detail shows only the count.
Model APIs
Retry advice spreads out under load
Congestion refusals now advise a randomized wait instead of a fixed 1 second, so refused clients do not all retry at the same instant.
  • 429 hosted_admission_congested and the 503s datastore_unavailable, hosted_evidence_unavailable, hosted_admission_unavailable and the connection gate’s limiter_unavailable send a Retry-After between 1 and 4 seconds, up to 8 seconds under heavy load.
  • The minimum stays 1 second. Codes, statuses and bodies are unchanged, so clients that honor Retry-After or Retry-After-Ms need no change.
See Errors.
Coding planBilling
Settings > Coding plan: usage limits and plan changes on one page

Settings > Coding plan

Settings > Coding plan is where you see and manage your coding plan. Billing keeps its coding plan section, with a Manage in Coding plan link, and existing links keep working.
  • Usage limits shows the 5-hour and weekly limits, Standby today and credits after limits as rows, with a bar against each limit (credits only with a cap). Each says when it resets, in your local time, such as Resets in 2 hr 13 min, Resets Sat 11:30 AM or Resets Oct 5. Standby Refills, and a row says Plan ends when your plan ends first.
  • An owner selects Subscribe to choose a plan. A member sees Ask an owner to choose a coding plan., and a workspace that cannot buy a plan sees the reason, with its next step where there is one.
  • The plan chooser marks the smallest tier that covers your last 30 days, without selecting it, and shows what the plan covers and what stays pay as you go.
  • After checkout, the page says after a minute that you can leave it. If the payment is still not confirmed after 10 minutes, it says Your payment is not confirmed yet. with Check again and Charged? Contact support.
  • Change plan reviews Plan, Starts, Due now and Then before you confirm: an upgrade starts now with a prorated charge, and a downgrade switches at renewal.
  • An unpaid upgrade shows Upgrade unpaid and a Pay button with the amount and the tier. If it is not paid by its deadline, your plan stays on its tier. Keep Pro, named for your current tier, takes back a scheduled downgrade in one press.
  • The cancel dialog leads with the day your plan ends, and its button says Cancel on that day, or Cancel now when cancellation is immediate. Your keys keep working with credits after the plan ends.
See Manage your plan.
CLI
CLI 0.3.3: a simpler full-screen app
The usual first run takes three presses of Enter. Start sign-in, continue from Agents, then approve Review.
  • Browser approval saves sign-in automatically. The next screen shows your name and workspace. Press Enter to save a pasted key. A newline that arrives with the paste never saves it.
  • A sign-in started while another RunInfra command is finishing waits briefly for it instead of failing.
  • Review starts with a model and a command name such as claude-run. Use m to change the model or e to rename the command. Some agents still need a model choice. Read Review to the end and clear its blockers before approving. Held or pasted keys never approve, and neither does a key pressed as Review appears or returns from Help or Diff.
  • Agents shows installed agents and the focused agent’s logo. Agents that need manual setup, or whose check failed, fold into one line that says which. Use m to show them or r to check again.
  • The payment question shows your credit balance alongside Pay as you go and Coding plan. It appears once per workspace when you can manage its plan, sales are open, and no plan exists. Escape goes back without choosing.
  • Footers show at most four hints. Press ? for Help and other shortcuts. Use Details for more information and Fix problems to check connections.
  • Done names the connected agent and the command to run. Home shows connected agents and, when data is available, requests, spend and cache savings for the last 24 hours.
  • runinfra disconnect --json reports notes such as a recovery snapshot in storageNotices, each with a level and message, instead of warnings, and counts restored files in one line.
See the CLI guide for the full-screen app and sign-in for browser and pasted-key setup.
Coding planAPI
Coding plans wait for running requests instead of stopping

Coding plans wait for running requests instead of stopping

When your running requests have reserved the rest of a plan limit, a new request now waits on the server for up to 30 seconds instead of moving to credits or stopping. It then runs on your plan as soon as a running request finishes. If none finishes in time, it gets a retryable 429 plan_window_busy. Only a limit that is really used up returns 402 plan_limit_reached.GET /v1/models documents max_concurrent_plan_requests_per_workspace and max_concurrent_standby_requests_per_workspace, the per-model allowances of a serving plan.
AccountSecurity
Settings > Security: sign-in methods, password and session controls, readable audit logs

Settings > Security

Settings > Security shows your email and linked sign-in methods. Password users can change their password after entering the current one. Google and GitHub users can request an email link to set a password. Sign out of other sessions ends every other session; existing access can last until its token expires.Account owners with audit logs see who made each change and when, with Access, Billing, Security and Account filters and a mobile event list. Switching from Personal to Business is free.
PricingCoding plan
Coding plan cards show requests at once per model; the calculator prices native FP8

Pricing and the cost calculator

Each coding plan card on Pricing states how many requests it can run at once per model, as a range when models differ.The inference cost calculator now prices checkpoints published in native FP8 through its cited path, and keeps withholding a price when the published precision does not match.
BillingAccount
Settings > Billing leads with your balance

Settings > Billing leads with your balance

Settings > Billing now opens on your balance, with auto-recharge directly under it. The coding plan, rate limits, partner pricing, spend limit, cost usage, invoices and billing contact follow in that order, when your account has them.Rate limits shows your three limits and your token tier. The per-model table and the token tiers are under All limits and tiers.Add funds, Manage billing details, auto-recharge, invoices and the billing contact are for the account owner. Members and viewers see the read-only sections, and their balance notes that an owner manages top-ups and invoices.
CLI
CLI 0.3.1 and 0.3.2: runinfra update and fixes
Run runinfra update to update the CLI.
  • npm and pip installations update through their own package manager. Standalone installations download the signed release, verify it and replace themselves.
  • Linux and macOS standalone installations stage updates in the background and apply them on the next start. Set RUNINFRA_NO_UPDATE=1 to turn this off.
  • The Windows standalone full-screen app now redraws correctly.
  • runinfra doctor keeps its directory explanations instead of hiding them.
See CLI updates for every channel and for recovery.
CLI
CLI 0.3.0: profiles and clearer setup
Run runinfra to connect coding agents from your terminal.
  • Relocatable agents get a separate launcher, such as claude-run, with their own settings.
  • Read Review to the end, then press Enter or y.
  • Re-login saves the new terminal key before revoking this machine’s previous key.
  • runinfra status shows live usage; --json returns a single report.
  • runinfra plan shows your coding plan. While plans are on sale, the first sign-in asks how you will pay.
Use the CLI guide for installation and JSON output for automation.
These headers are not live yet. This entry has no ship date. Until they ship, read usage.prompt_tokens_details.cached_tokens in the response body.

A chat completion reports what the session cache did for it

When released, a non-streaming POST /v1/chat/completions response to a request that resolved a session identity will carry three session headers. You resolve one by naming a session with the identity headers documented in Session affinity: x-session-id, x-parent-session-id, or x-session-affinity. A request that sends no identity header and no prompt_cache_key may still resolve a session; Session affinity states when it does.
  • x-session-cached-share is the figure to read: the cached input tokens as a share of this request’s prompt tokens, 0.00 to 1.00 with two decimals, and 0.00 for a request with no input tokens. It is the share of your input billed at the cached input rate.
  • x-session-state reports the session’s cache state: cold, warm, or restored.
  • x-session-tier names where the session’s prefix is held, gpu, ram, or nvme. ram is the tier GET /v1/models lists as host_ram.
Which responses will carry them:
  • A non-streaming response to a request that resolved a session identity carries all three.
  • A streamed response carries no session headers.
  • A request routed by prompt_cache_key alone or kept at workspace-level placement carries none, and a /v1/messages call carries none.
  • They appear on every hosted model whose GET /v1/models entry reports cache_isolation as isolated. How a session is placed and how long its prefix is kept is in Session affinity. GET /v1/models states each model’s caching behavior as cache_isolation, cache_tiers, and cache_retention.
APIModel APIs
Coding plans for coding agents

Coding plans for coding agents

Starter ($10), Pro ($29) and Team ($99) monthly plans pay first for covered Model API requests from your workspace keys, with a 5-hour limit and a weekly limit. At a limit, credits pay if you chose them, up to your cap, then free Standby serves chat. See Coding plan for tiers and coverage, and Coding plans in the CLI to connect your agents and read your plan from the terminal.
APIModel APIs
A paused model gives the same Retry-After on every endpoint

A paused model gives the same Retry-After on every endpoint

503 hosted_model_paused now carries the same Retry-After on every endpoint: the time until the next availability check, capped at one hour, and 60 seconds once that check time has passed. Chat completions, /v1/messages and /v1/responses already worked this way. On /v1/embeddings and the other non-chat endpoints (audio, images), a pause whose check time had passed could ask for a retry after 1 second. See Errors.
APIModel APIsSecurity
Copy an API key again instead of making a new one

Copy a key again instead of making a new one

Settings > API keys, the model pages and the dashboard home now offer Copy key. The member who created a key and the workspace owner can copy it, and every copy is recorded in the audit log. A key made before September 25 exists only as a one-way digest: rotate it once to get a copyable replacement. See Authentication.
APIModel APIs
Coding agents recover from a full context window, keyed retries are protected for the whole request, and more requests are accepted as sent

Coding agents recover from a full context window

A prompt longer than the model’s context window now returns each API’s own overflow error, so coding agents compact and continue instead of resending. Chat completions return 400 with code context_length_exceeded and error.context_window. /v1/messages returns prompt is too long, which Claude Code compacts on. A streaming /v1/responses request receives a response.failed event with code context_length_exceeded, which Codex compacts on before its next turn. Nothing is charged for the refusal. See Errors.

Idempotency-Key covers the whole request

A request with an Idempotency-Key now holds the key for as long as it can run: 830 seconds at most, streamed or not. A retry sent while a slow non-streaming request is still generating gets 409 idempotency_conflict instead of starting a second generation. When a key cannot be checked, a keyed request on every endpoint now returns 503 idempotency_unavailable with Retry-After before any work or charge; on /v1/messages it arrives as overloaded_error. Requests without the header are unchanged. See Idempotent retries.

Requests accepted as sent

Qwen3.8 27B now sends reasoning_effort: "max" as xhigh, its maximum and default, instead of refusing it, as it already did for high. Chat completions also accept null for the optional fields OpenAI documents as nullable, treating it as absent, and text with an unpaired UTF-16 surrogate (a truncated emoji, for example) is repaired to U+FFFD instead of failing the request.
APIModel APIs
Coding agent sessions are recognized from the headers the agents already send

Coding agents keep their sessions together

Claude Code, Codex CLI, and OpenCode sessions are now recognized from the session headers these agents already send. Each conversation keeps its requests on the same warm cache with nothing to configure, and each Claude Code subagent is a session of its own. A request’s own x-session-affinity now wins over x-parent-session-id. OpenCode sends its own session in x-session-id as well, so it already kept each subagent on its own session. A malformed session header no longer hides a valid one sent with it: the next header in the order is read instead. See Session affinity.
APIModel APIsModels
Claude Code and Codex CLI sessions run on the Qwen3.8 models and Ornith 1.5 35B

Coding agents on the Qwen3.8 models and Ornith

Claude Code and Codex CLI sessions now run on Qwen3.8 27B, Qwen3.8 2.4T A95B, Qwen3.8 Flash Next, and Ornith 1.5 35B. These models accept one system message, and only as the first message. Both agents send more than one: Claude Code adds system messages during the conversation, and Codex sends its instructions together with a developer message. The gateway now merges leading system messages and sends later ones as user messages in the same position. Requests these models used to refuse with invalid_message_order now succeed, and each turn still extends the previous prompt, so prompt caching keeps working. See System message placement.Codex CLI needs web_search = "disabled" in config.toml, because its hosted web search tool runs on OpenAI’s servers. See Codex.
APIModel APIsCredits
Credits responses name the workspace behind the API key

Identify the workspace behind a credits response

GET /v1/credits now includes workspace_id, the UUID of the workspace the API key belongs to. You can read credits only with a workspace API key. See Credits and budget.
AccountBilling
Business invoice details on the account page and one billing contact for billing notices

Invoice details on the account page

A Business account now shows the company details its invoices carry at Settings > Account. The company name, billing address, and tax ID are read-only there. Select Edit in billing to open the billing portal, or set them at checkout.

One billing contact for billing notices

An owner can enter one email address as the billing contact at Settings > Billing. Purchase receipts, balance and spend alerts, payment issues, plan changes and account holds go here, with the owner copied. The invoice for a purchase follows the payment account, and this setting cannot move it. Clear the address to send those notices to the owner alone. The billing contact gets no account access, sign-in, or permissions. Security and account emails, such as sign-in links and API key expiry warnings, always go to the owner and never to the billing contact.

Rate limits that say what they are

Settings > Billing now names the three limits on your account. Requests per minute are per API key. Tokens per minute and concurrent requests are per model. Concurrency shares apply only while a model is busy. With spare capacity, your account can exceed its share. A per-model table shows the current numbers. The token-budget ladder shows Model defaults, At least 20M from 50inlifetimepurchases,Nopracticallimitfrom50 in lifetime purchases, No practical limit from 200, and Enterprise by arrangement. A model keeps a higher default. Pricing and credits explains how to raise each limit.

Partner pricing

Your account’s agreed Model API prices now appear while you are signed in. They apply automatically to requests. RunInfra arranges prices per account and per model. Partner pricing at Settings > Billing lists each covered model with your price beside the list price. It appears only for accounts with arranged prices. GET /v1/models returns your account’s prices for its API key. The response’s usage.cost reflects your agreed price. Partner pricing is not self-serve. Contact us.
AccountTeamsSecurityAPIModel APIs
Business accounts, 14-day invitations, account-wide audit logs, and cached-input usage

Business accounts

Only an account owner can switch between Personal and Business. Open Settings > Account and set Account type. A Business account uses the company name in Account name. The switch is reversible.Invoices carry the company details you provide. Set the company name, billing address, and tax ID at checkout or under Manage billing details on Settings > Cost.

Invite people before they sign up

Settings > Members now accepts an email that does not yet have a RunInfra account. If the address already belongs to a RunInfra account, that person gets access right away. Otherwise, the invite is valid for 14 days, and the person joins the inviting account as soon as they sign up with that address. Until then, the owner sees a Pending row and can select Revoke.

Audit logs for Business accounts

Audit logs are available on Business accounts and Enterprise plans at Settings > Security. A Business account’s owner can review every action recorded in the account.

The cached input count on every usage object

usage.prompt_tokens_details.cached_tokens is now present on every hosted model response, together with the same number as usage.runinfra.cached_input_tokens beside cost_microcents. It is the count of input tokens billed at the cached input rate, the figure the request’s cost was computed with, so on a billable response prompt_tokens - cached_tokens is exactly what was billed at the uncached rate. It is 0 rather than absent when nothing was billed as cached, so a client never has to interpret a missing field; that includes a response that settled at zero and a model whose cache is shared across tenants, where hits are neither disclosed nor priced. On a stream both fields ride the usage frame sent with [DONE], beside the cost; a /v1/responses usage carries the count as input_tokens_details.cached_tokens and runinfra.cached_input_tokens; the idempotent replay of a streamed request carries the figure that was settled. The contract is in Chat completions.
Model APIsDashboardAPI
Best-account cache hit rates and concurrency allowances that apply when a model is busy

The cache hit rate on every model is the best account’s rate

Every surface that states a model’s cache hit rate now states the best account’s rate, not an average across all accounts. The figure is the highest share of billed input tokens served from the cache that any one account reached on that model over the trailing day, measured on multi-turn agent sessions from accounts that sent at least ten million billed input tokens in that window, and it is withheld for a model whose best account sits under 70 percent. It appears with the words “best account” beside it on the model page, in the Model Library, on Pricing, and on the model’s dashboard page, and the effective input rate line derives from the same figure. An average across accounts was dominated by whichever account sent the most traffic, so it understated what a session that stays on the model actually gets. Your own workspace’s rate is on the usage page and on the dashboard home, unchanged.

Spare capacity is open to you

A key or workspace is no longer held to its concurrency allowance while the model has capacity to spare. The allowance published as max_concurrent_requests_per_api_key is what each workspace is held to once a model is busy, which keeps a busy model shared rather than taken by whoever arrived first. While the model has capacity to spare, the gateway admits your requests past it, up to a higher ceiling, and the allowance applies again as the model fills. A refusal looks exactly as before: hosted_per_key_concurrency_limit and hosted_workspace_concurrency_limit keep their envelope and headers, and error.limit quotes the limit that applied at that moment. Details in Rate limits.Paying workspaces are held to a larger share. Any workspace that has bought credit, at any amount, is held to the same share of a busy model as a provider partner, which is most of the model. Read your allowance from max_concurrent_requests_per_api_key in GET /v1/models.
APIModel APIsModels
Qwen3.8 Flash Next serves its full 1,048,576 token window and accepts image input

Qwen3.8 Flash Next: the full window and image input

Qwen3.8 Flash Next now serves a 1,048,576 token context window, up from 262,144, at the same prices. Retrieval was verified at 10, 50 and 90 percent depth of prompts over one million tokens before the number changed, and the model’s answers on a fixed short-prompt set were unchanged. GET /v1/models reports the new context_window, and a prompt above it is refused with 400 naming the limit, never truncated.Qwen3.8 Flash Next accepts image input. Send image_url or input_image content parts as inline data URLs (png, jpeg, webp or gif), up to 8 images per request. Image tokens are counted by the server inside prompt_tokens and bill at the model’s input rate; there is no separate image charge. Remote image URLs are not fetched. The contract is in Chat completions.
APIModel APIsAccountDashboard
Qwen3.8 Flash Next joins the hosted models, an optional monthly spending limit per API key, per-key model usage, and a cache-hit rate on every model

Qwen3.8 Flash Next joins the hosted models

Qwen3.8 Flash Next is live. Qwen3.8 Flash Next, served at its vendor-released FP8 precision, with a 262,144 token context window and a 262,144 token output budget, priced at $0.12 per million input tokens, $0.01 per million cached input tokens, and $0.40 per million output tokens. It calls tools, returns JSON mode, streams, and caches prefixes, so a repeated prompt prefix bills at the cached rate. Reasoning is on by default at the model’s highest effort; reasoning_effort accepts low, medium, high, and xhigh, and none turns it off. Its page in the Model Library carries the full published facts.

A spending limit on one key

Any API key can carry an optional monthly spending limit. Set it when you create the key or later from its row under Settings > API keys, and raise or remove it at any time. Once the key’s settled spend in the current calendar month (UTC) reaches the limit, that key’s requests are refused with 402 spend_limit_reached carrying limit_cents, spent_cents, and resets_at; the workspace’s other keys and its credits are unaffected, and the month resets on the first. The contract is in Errors and Authentication.

Usage by key, by model

The usage page’s By API key table opens each key onto the models it called, with that pair’s own cost, requests, tokens, and cache measurement, so a router key shows where its spend went. Each model row now leads with the model’s logo.

A cache-hit rate on every model row

The Model Library states the measured cache-hit rate on every model, over the trailing day, with the window named beside the figure.
APIModel APIsClaude Code
Anthropic-compatible tools, images, thinking, token counting, and Claude Code support

Anthropic-compatible Messages supports agent workflows

POST /v1/messages now supports client tools, base64 image blocks on compatible models, thinking controls, structured output, and Anthropic event-named streaming. Tool results, prior thinking blocks, consecutive same-role turns, usage fields, request ids, and Anthropic error envelopes follow the contract in Anthropic Messages.POST /v1/messages/count_tokens estimates the input size without running or billing the model. It accepts the same request body, with max_tokens optional, and uses the same authentication and per-key rate limit as Messages.Claude Code can use RunInfra hosted models through the Messages API. The new Claude Code guide covers the base URL, authentication variables, model aliases, subagents, streaming, thinking, and usage checks.
APIModel APIsCreditsAccount
a credits endpoint, the exact cost of a request on its usage object, image input on GLM 5.3 Flash, and a 404 that says a model retired

Read your balance without earning a refusal

GET /v1/credits returns your workspace balance. A workspace key reads balance, held, available, period spend, spend cap, and plan tier, so a client can see where it stands before a request earns a 402 instead of learning it inside the refusal. The operation is read-only and uncached, and carries the same envelope and rate-limit headers as GET /v1/models. A cap that is not configured is reported as null rather than as a number.

Every request reports what it cost

usage.cost states the request’s charge in US dollars, computed by the same settlement formula the ledger uses, with the exact ledger unit beside it in usage.runinfra.cost_microcents. A stream prints it on the usage-bearing frame at the end, and a stream with no billable output prints zero. It reaches /v1/responses and idempotent replays. Hosted models only.

GLM 5.3 Flash accepts image input

GLM 5.3 Flash takes images as inline data URLs, proven with multi-image and concurrent requests. Remote image URLs are not fetched. There is no separate image charge.

A retired model’s 404 says it retired

A call naming a retired model still returns 404 model_not_found, and the body now carries the retirement date and a link to the Model Library. The status and the code are unchanged on purpose, so a handler keyed on them keeps working.

Sign-in

An expired confirmation link can be resent from the sign-in page. The email step opens on the expired link and offers a fresh confirmation email under the address field.
Model APIsModelsAPIBillingAccount
GLM 5.3 Flash is public, four models retire, the legacy tool dialect is refused, and team accounts arrive

GLM 5.3 Flash is public

GLM 5.3 Flash joins the hosted models. A 1,048,576 token context window and a 32,768 token output budget, priced at $0.10 per million input tokens, $0.01 per million cached input tokens, and $0.40 per million output tokens. It calls tools, and it caches prefixes, so a repeated prompt prefix bills at the cached rate. Its page in the Model Library carries the full published facts.

Speech to text, embeddings, and rerank are retired

Parakeet TDT 0.6B v3, Qwen3 Embedding 0.6B, Qwen3 Embedding 8B, and Qwen3 Reranker 8B no longer serve. A call naming any of them returns 404 model_not_found. No hosted model answers POST /v1/audio/transcriptions, POST /v1/embeddings, or POST /v1/rerank today. GET /v1/models is the authoritative live list.A link to a removed model lands on the Model Library rather than on a dead page.

The legacy tool dialect is refused

functions and function_call return 400 on DeepSeek V4 Flash. The deprecated fields were previously accepted and silently discarded, so the definitions never reached the model and the reply came back as prose no tool parser could read. tools and tool_choice are the supported dialect. The refusal is declared per model, and DeepSeek V4 Flash is the model it was measured on.

Nemotron 3.5 Lightning earns its cached rate

nemotron-3-5-lightning-30b bills cached input at $0.01 per million tokens. A cached rate is published only where prefix-cache isolation has been verified on the served model. Nemotron’s attestation landed, so the rate its row carried is now the rate you are charged.

Team accounts

You can invite members, switch accounts, and leave an account. Members is a settings destination with two roles, owner and member. Billing and the commercial relationship stay with the owner.
APIModel APIs
a session header that keeps a conversation on the same warm prefix cache
x-session-id and x-parent-session-id are first-class session identities. Precedence is x-session-id, then x-parent-session-id, then x-session-affinity, so a subagent stream carrying only its parent’s id shares the parent’s warm prefix cache. A request that sends none of the three is byte identical to before.
APIModel APIsReliability
longer response ceilings, a larger share of a model's concurrency, a free tier that is not one request per second, and large idempotent replays
Response time ceilings went up. Maximum request duration moved from 300 to 800 seconds, and the platform response ceiling from 240 to 740 seconds, after a customer running long non-streaming completions failed 9 of 10 calls at the old ceiling. The time allowed for a streaming response to begin moved from 60 to 180 seconds; a non-streaming response has the full ceiling.A paying workspace gets half of a model’s concurrency slots, not a quarter. Never the whole model, so a second paying tenant can still get in.The free tier’s requests per minute moved from 60 to 10,000. One request per second is below what a single agentic client idles at. The published tier table now matches what is enforced.A large response can be idempotently replayed. The replay record ceiling went from 512 KiB to 6 MiB. Above the old ceiling, a retry with the same Idempotency-Key used to get a permanent 422.
BillingAPIModel APIs
spend raises the tokens-per-minute budget, and a token meter that had been counting bytes
Spend tiers raise the tokens-per-minute budget. Starter keeps the model default, unchanged for every existing workspace, and Scale and Pro raise it. Qualification is by cumulative amount paid. Concurrency is deliberately not tiered: the token meter is a fairness budget, and the concurrency share is the capacity guard. The dashboard draws all three tiers with your current one marked.The token meter stopped counting bytes. It was measuring bytes rather than tokens, so every workspace’s per-minute budget was effectively a quarter of the published figure.
Model APIsModelsAPI
Ornith 1.5 35B is public, and every refusal says when to retry

Ornith 1.5 35B is public

ornith-1-5-35b joins the hosted models. A 262,144 token context window and a 32,768 token output budget, priced at $0.10 per million input tokens, $0.01 per million cached input tokens, and $0.40 per million output tokens. It accepts text and image input, and it caches prefixes.

Every refusal says when to retry

A paused model and a rate limit both carry Retry-After. A paused refusal sets it to the distance to the next availability check. A 429 sets it from the state of the limiter that refused you, in seconds and in milliseconds. A capability or parameter refusal carries no retry header, because the same request will not succeed later.
Model APIsModelsBilling
DeepSeek V4 Pro is public, image input on Qwen3.8 27B, one hosted checkout, and a measured cache hit rate on model pages

DeepSeek V4 Pro is public

DeepSeek V4 Pro joins the hosted models. A 1,048,576 token context window and a 32,768 token output budget, priced at $0.60 per million input tokens, $0.03 per million cached input tokens, and $1.90 per million output tokens. Text input, and it caches prefixes.

Image input on Qwen3.8 27B

qwen3-8-27b accepts images as inline data URLs. Remote URLs are not fetched. Accepted input is now a per-model fact, and it drives each model page’s Accepted input row, the refusal a text-only model returns, and the input modalities we publish to catalogs.

One checkout

Every purchase goes through hosted checkout. The purchase carries a durable identifier for that attempt, returns to the page you started from, verifies the grant on return, and survives a back button. Billing details are collected at checkout.

Model pages publish a measured cache hit rate

A model with a cached input price publishes the share of input tokens actually served from cache. It is derived from settled billing over a trailing day, and it is absent rather than estimated when the measurement cannot be made.
APIPrivacyModel APIs
model discovery publishes context and prices, and a retention statement that matches what is stored
GET /v1/models and GET /v1/models/{model} publish context_length and a pricing object. Rates are in US dollars per token, and a cached rate appears only for a model whose prefix-cache isolation is attested. A newly registered model publishes both with no code change.The data retention statement is exact. The policy said content is never stored, while the idempotent replay cache holds non-streaming response bodies for 24 hours by design. It now states that window and what it is bound to, notes that streaming is exempt and that a keyless request leaves no stored content, and every model page carries a Data retention row from the same source these docs read. See Data retention.
APIBillingCreditsModel APIs
Idempotency-Key covers streaming, cached input pricing, credit top-ups that cannot shrink or double charge, and a named milestone when you cross one

Idempotency-Key now covers streaming

A dropped stream retried with the same Idempotency-Key no longer generates or charges twice. Streaming chat completions used to ignore the header. They now record the key before the model runs, so a duplicate arriving while the stream is open gets 409 idempotency_conflict and starts no second generation, and a duplicate arriving after the original settled gets the terminal usage and cost as JSON with X-RunInfra-Idempotent-Replay: true. We do not store the tokens we deliver, so the text itself is never replayed. A streaming key is held for the response deadline plus a short settlement grace, 270 seconds at the outside, then clears on its own. The bound follows the response ceiling, so since the August 22, 2026 change it is 770 seconds.This changes behavior for a client that sends one constant key: give each logical request its own key, and reuse a key only to retry that same request. The contract is in Idempotent retries.

Cached input is priced separately

A repeated prompt prefix bills at a cached input rate, on the models where cache isolation is proven. On an attested model, cache isolation is per workspace: a prefix cached for one workspace is never served to another. A model earns a published cached rate only once that isolation has been verified on the served model, with an identical prefix under a different workspace’s cache identity confirmed to return no cached tokens. DeepSeek V4 Flash, qwen3-8-27b, and Qwen3.8 2.4T A95B are attested today. A model with no attestation publishes no cached price and bills cached input at its standard input rate.

Credit top-ups

Promotion codes no longer apply to credit top-ups. Credits are stored value, so a discount code reduced the credits you received by exactly what it took off the price: it appeared to work and bought you nothing. The amount you select is now the amount you are charged and the amount that lands on your balance.A reload cannot charge you twice. A top-up carries a durable identifier for that attempt, so a refresh, a back button, a second submit, or a return from a bank verification screen resumes the same payment instead of starting another.A confirmation that does not settle is reported as unknown, never as declined. When confirmation does not finish inside its bound, the screen says the result is not yet known and waits for the credits to land, rather than reporting a failure on a payment that may have succeeded.A captured payment now ends in credits or a retry. A top-up settles only when the amount quoted, the amount the payment provider recorded, and the amount actually captured all agree in US dollars.A 402 hands an API caller an absolute top-up URL, so the refusal points straight at the page that clears it.

Crossing a spend milestone says so

The purchase that moves you up a capability milestone names the milestone you reached and the one you left. It reads the same ladder that enforces your limits, so it cannot congratulate you on capacity your workspace has not been granted, and it only moves forward.

Model pages state how a number was measured

A model’s throughput figure now names its measurement path in the unit, reading output tokens per second, end to end when measured through the API you call, and output tokens per second, model only when measured at the model itself. The two are different measurements, and printing them under one unit let a model-only figure read as the speed a request receives.
Model APIsModelsAPI
Model APIs go live: hosted models behind one OpenAI-compatible endpoint, a second model, and reasoning-model budgets

Model APIs are live

RunInfra now hosts and serves models behind one OpenAI-compatible API. Point an OpenAI client at https://api.runinfra.ai/v1, create a workspace key, and call the model id. No GPU to provision and nothing to deploy, priced per million tokens at the rates each model’s page publishes. Start at the Model APIs quickstart.The Model Library is the public catalog. Every hosted model’s page publishes its context window as served by our own deployment, its prices, its capabilities, and its measured performance with the conditions of the measurement attached.Nemotron 3.5 Lightning 30B (nemotron-3-5-lightning-30b) joins DeepSeek V4 Flash, listed ahead of its public opening with its served facts published. A paused model answers 503 hosted_model_paused with its scheduled return time and stays listed in GET /v1/models, so nothing about your integration changes when it opens.

Reasoning models budget correctly out of the box

Reasoning models spend billed output tokens thinking before they answer, so an under-budgeted request can return empty content. Each reasoning model’s page now publishes its recommended minimum max_tokens and bakes it into the code examples it generates. See Reasoning models.
BillingAccountSecurity
one honest credit balance, self-serve account deletion, and a hardened sign-in
One balance, everywhere. The navbar and billing page read the same live available balance, updating after credit purchases and settled usage, so every surface reports the same amount.Signup credits are visible from day one. The grant and every transaction against it appear on Settings > Cost, and the live balance shows in the top-bar credits chip on every page.Self-serve account deletion. Settings > Workspace has a danger zone. The dialog previews exactly what will be removed, lists any blockers with inline actions to clear them, and requires a typed confirmation plus your password.Sign-in hardening. Passwords require 8 characters minimum. Email confirmation and password-reset links work when opened in a different browser, and expired links say so with a resend path.
APIDocs
Responses adapter reference, and an expanded error reference
Responses adapter. A dedicated /v1/responses reference for the Responses-shaped chat-completions adapter, including streaming, instructions, response_format, and supported tool pass-through fields.Error reference. The OpenAPI spec and error guide now cover the public gateway statuses developers should handle: rate limits, credit exhaustion, idempotency conflicts, replay-unavailable responses, upstream failures, and gateway timeouts. See Errors.