Skip to main content
This endpoint reads the workspace credit balance and budget for pay-as-you-go work. A coding plan can fund covered requests without spending that balance. Use GET /v1/usage for coding plan windows and policy; /v1/credits does not include a coding_plan field. It is read-only and charges nothing. It is authenticated and rate-limited like every other endpoint, and the response is never served from an HTTP cache.

Request

The OpenAI SDKs have no credits method, so call the endpoint with your HTTP client. The same key you use for inference works here, as long as it is a workspace key (see Who can read it).

Response

Every amount is an integer in US cents. Nothing here is a dollar float, because the ledger stores cents and a float would be the first rounding in the chain.

What decides whether a request runs

A credit-funded inference request is admitted against available_cents. Holds for requests already in flight are already taken out of that number, so a long request on a thin balance can show a small available_cents and a positive held_cents at the same time; when the request settles, the unused part of its hold returns to the balance. The legacy workspace spend cap does not enter this decision, which is why gates_inference is always false. A per-key spending limit is separate and can refuse requests on that key. A request refused for insufficient credits answers 402 with code: "insufficient_credits" and error.current_balance_cents, the same balance this endpoint reports, plus a topup_url. See Errors. The credit amounts above appear directly in error on OpenAI-compatible refusals. The Anthropic Messages error envelope preserves only the error type and message text.

Who can read it

Use a workspace API key created in Settings. A key that can spend the balance can read it, because the number is already disclosed to that key inside every 402 it earns. Adding funds, changing the plan and reading invoices stay with the account owner in the dashboard; there is no write surface here.

Freshness and limits

The response carries Cache-Control: no-store and is read from the ledger on every call. Use as_of rather than your own clock when you compare two reads. On top of your workspace’s per-key rate limit this endpoint has its own allowance of 60 reads a minute per key. Poll it when you need a decision, not on a timer. The x-ratelimit-* headers on the response describe the plan bucket, so a boot-time read here also tells you your remaining request quota. During a rate-limit store outage, both limits are counted per server.

Errors