POST /v1/chat/completions at https://api.runinfra.ai/v1 with a workspace API key.
Using a coding agent? Your agent can set up RunInfra’s open models for you. Copy the prompt from Set up with your coding agent.
1
Get a workspace API key
Open the Model Library, pick a model that is not marked Paused, and select Get API key. If you already have a key you can copy, the page highlights that key’s copy button instead of making another one.
2
Put the key in your environment
3
Make the call
What comes back
What comes back
A normal completion. Token counts and ids are illustrative; Check
reasoning carries the model’s thinking, and content carries the answer.finish_reason before you use content. stop means the model finished on its own. length means it hit the output budget. On a reasoning model, that can leave content empty with the budget spent on thinking, which is why these samples set max_tokens to 16,384.The live chat model ids
RunGET /v1/models for the current list; a model that is temporarily paused is reported there and on its Model Library page.
How usage is counted
The token counts in the responseusage object are the billing counts. For chat completions, cached input is reported at usage.prompt_tokens_details.cached_tokens on models that publish a cached-input price, when a cached count is available for the request. Those tokens remain part of prompt_tokens and are billed at the model’s cached-input rate.
A client-side tokenizer can disagree with the billed count because it may use a different tokenizer revision or omit the exact chat template, special tokens, and media preprocessing that are applied when the request is served. Use the response usage fields and the usage dashboard for reconciliation.
Why max_tokens is 16,384
These models reason before they answer, and reasoning tokens are billed output that count towardmax_tokens.
You pay for what is generated, not for the ceiling you allow, so a generous ceiling costs nothing and a small one can cost you a whole request.
Next steps
Chat completions
Every field the gateway forwards, validates, or drops.
Streaming
Read deltas and ask for the usage frame.
Models
Discover the ids your key can call.
Coding agents
Paste one prompt. Your coding agent sets up RunInfra’s open models for you.