Skip to main content
This is how an application discovers, at boot, which model ids its key can call and what binds them. Never hard-code a list you guessed.

Request

Response

The numeric values above are an example. Read them from your own response, because they move with capacity. The example uses Nemotron 3.5 Lightning 30B’s published 262,144-token context and maximum generated output. The fields are additive.

The limit fields exist so you never learn a limit from a 429

They are read from the code that enforces them, so a published limit and an applied limit cannot drift apart. All of them are additive to and disjoint from the fields the OpenAI Model object declares, so a client typed against the OpenAI SDK keeps parsing this response and ignores what it does not recognize.

Absent means undeclared, never zero and never unlimited

A field appears only where the limit it names is actually enforced for that model. A model that has not reported a served window carries no context_window rather than borrowing another model’s number.

Bytes can bind before tokens

max_request_bytes is a byte ceiling, not a token ceiling, and the two bind independently. For a very long prompt the byte ceiling is often the one that binds first, whatever the context window says.

Prices and cache behavior

Hosted model prices are returned here for discovery and repeated on each model’s page in the Model Library. The page also carries the human-readable cache eviction and retention statement, served precision when verified, capabilities, and the recommended output budget.

Available and paused

Model availabilityGET /v1/modelsavailableavailability: available200requests are served and billedListed with its id, prices on its page.pausedavailability: paused503hosted_model_pausedNothing is charged. Honor retry headers.pausereturnA paused model is not a missing model: the id stays valid and stays listed, with paused_untilcarrying the next availability check. Poll and retry; never drop the id from your configuration.
Model availabilityGET /v1/modelsavailableavailability: available200requests are served and billedListed with its id, prices on its page.pausedavailability: paused503hosted_model_pausedNothing is charged. Honor retry headers.pausereturnA paused model is not a missing model: the id stays valid and stays listed, with paused_untilcarrying the next availability check. Poll and retry; never drop the id from your configuration.
A paused model keeps its place in the list. It is a real configured model that is temporarily not serving, so poll this endpoint or retry the call itself to detect it coming back. The id does not change across a pause. While paused, the row omits pricing, the cached-input price fields, max_concurrent_requests_per_api_key and max_tokens_per_minute_per_workspace.

Retrieve one model

Returns a single object with the same fields and no list wrapper. An id your key cannot reach returns 404 model_not_found. URL-encode any id containing a slash.

Which models appear

A workspace key lists every public hosted model, available or paused.

Authentication and rate limits

The endpoint needs a workspace API key, sent as Authorization: Bearer or x-api-key. It uses that key’s requests-per-minute limit and returns X-RateLimit-Limit, X-RateLimit-Remaining, X-RateLimit-Reset and X-RateLimit-Tier. Listing models reserves no inference credit, but it is still an authenticated request against your per-minute budget, so do not call it in a hot loop.
Use the returned id exactly as it appears. It is the supported name for the model field.

Chat completions

Send one of these ids as the model field.

Authentication

The key that decides which ids you see.

Rate limits

The four layers that can refuse you.