Request
Response
The fields are additive.
The limit fields exist so you never learn a limit from a 429
They are read from the code that enforces them, so a published limit and an applied limit cannot drift apart. All of them are additive to and disjoint from the fields the OpenAI Model object declares, so a client typed against the OpenAI SDK keeps parsing this response and ignores what it does not recognize.Absent means undeclared, never zero and never unlimited
A field appears only where the limit it names is actually enforced for that model. A model that has not reported a served window carries nocontext_window rather than borrowing another model’s number.
Bytes can bind before tokens
max_request_bytes is a byte ceiling, not a token ceiling, and the two bind independently. For a very long prompt the byte ceiling is often the one that binds first, whatever the context window says.
Prices and cache behavior
Hosted model prices are returned here for discovery and repeated on each model’s page in the Model Library. The page also carries the human-readable cache eviction and retention statement, served precision when verified, capabilities, and the recommended output budget.Available and paused
A paused model keeps its place in the list. It is a real configured model that is temporarily not serving, so poll this endpoint or retry the call itself to detect it coming back. The id does not change across a pause. While paused, the row omitspricing, the cached-input price fields, max_concurrent_requests_per_api_key and max_tokens_per_minute_per_workspace.
Retrieve one model
list wrapper. An id your key cannot reach returns 404 model_not_found. URL-encode any id containing a slash.
Which models appear
A workspace key lists every public hosted model, available or paused.Authentication and rate limits
The endpoint needs a workspace API key, sent asAuthorization: Bearer or x-api-key. It uses that key’s requests-per-minute limit and returns X-RateLimit-Limit, X-RateLimit-Remaining, X-RateLimit-Reset and X-RateLimit-Tier. Listing models reserves no inference credit, but it is still an authenticated request against your per-minute budget, so do not call it in a hot loop.
Use the returned
id exactly as it appears. It is the supported name for the model field.Related
Chat completions
Send one of these ids as the
model field.Authentication
The key that decides which ids you see.
Rate limits
The four layers that can refuse you.