Skip to main content
This is how an application discovers, at boot, which model ids its key can call and what binds them. Never hard-code a list you guessed.

Request

Response

An abbreviated example. Read the values from your own response; they move with capacity.

Notes

  • The limit fields exist so you never learn a limit from a 429. They are read from the code that enforces them, so a published limit and an applied limit cannot drift apart.
  • These fields are additive to and disjoint from the fields the OpenAI Model object declares, so a client typed against the OpenAI SDK keeps parsing this response and ignores what it does not recognize.
  • Absent means undeclared, never zero and never unlimited. A field appears only where the limit it names is actually enforced for that model; a model that has not reported a served window carries no context_window rather than borrowing another model’s number.
  • max_request_bytes is a byte ceiling, not a token ceiling, and the two bind independently. For a very long prompt the byte ceiling is often the one that binds first, whatever the context window says.
  • Hosted model prices are returned here for discovery and repeated on each model’s page in the Model Library, with the human-readable cache eviction and retention statement, served precision when verified, capabilities, and the recommended output budget.
  • A workspace key lists every public hosted model, available or paused.

Available and paused

Model availabilityGET /v1/modelsavailableavailability: available200requests are served and billedListed with its id, prices on its page.pausedavailability: paused503hosted_model_pausedNothing is charged. Honor retry headers.pausereturnA paused model is not a missing model: the id stays valid and stays listed, with paused_untilcarrying the next availability check. Poll and retry; never drop the id from your configuration.
Model availabilityGET /v1/modelsavailableavailability: available200requests are served and billedListed with its id, prices on its page.pausedavailability: paused503hosted_model_pausedNothing is charged. Honor retry headers.pausereturnA paused model is not a missing model: the id stays valid and stays listed, with paused_untilcarrying the next availability check. Poll and retry; never drop the id from your configuration.
A paused model keeps its place in the list. It is a real configured model that is temporarily not serving, so poll this endpoint or retry the call itself to detect it coming back. The id does not change across a pause. While paused, the row omits pricing, the cached-input price fields, max_concurrent_requests_per_api_key and max_tokens_per_minute_per_workspace.

Retrieve one model

Returns a single object with the same fields and no list wrapper. An id your key cannot reach returns 404 model_not_found. URL-encode any id containing a slash.

Authentication and rate limits

The endpoint needs a workspace API key, sent as Authorization: Bearer or x-api-key. It uses that key’s requests-per-minute limit and returns X-RateLimit-Limit, X-RateLimit-Remaining, X-RateLimit-Reset and X-RateLimit-Tier. Listing models reserves no inference credit, but it is still an authenticated request against your per-minute budget, so do not call it in a hot loop.

Chat completions

Send one of these ids as the model field.

Authentication

The key that decides which ids you see.

Rate limits

The four layers that can refuse you.