/v1/credits does not include a coding_plan field.
It is read-only and charges nothing. It is authenticated and rate-limited like every other endpoint, and the response is never served from an HTTP cache.
Request
Response
Every amount is an integer in US cents. Nothing here is a dollar float, because the ledger stores cents and a float would be the first rounding in the chain.
What decides whether a request runs
A credit-funded inference request is admitted againstavailable_cents. Holds for requests already in flight are already taken out of that number, so a long request on a thin balance can show a small available_cents and a positive held_cents at the same time; when the request settles, the unused part of its hold returns to the balance.
The legacy workspace spend cap does not enter this decision, which is why gates_inference is always false. A per-key spending limit is separate and can refuse requests on that key.
A request refused for insufficient credits answers 402 with code: "insufficient_credits" and error.current_balance_cents, the same balance this endpoint reports, plus a topup_url. See Errors.
The credit amounts above appear directly in error on OpenAI-compatible refusals. The Anthropic Messages error envelope preserves only the error type and message text.
Who can read it
Use a workspace API key created in Settings. A key that can spend the balance can read it, because the number is already disclosed to that key inside every402 it earns.
Adding funds, changing the plan and reading invoices stay with the account owner in the dashboard; there is no write surface here.
Freshness and limits
The response carriesCache-Control: no-store and is read from the ledger on every call. Use as_of rather than your own clock when you compare two reads.
On top of your workspace’s per-key rate limit this endpoint has its own allowance of 60 reads a minute per key. Poll it when you need a decision, not on a timer. The x-ratelimit-* headers on the response describe the plan bucket, so a boot-time read here also tells you your remaining request quota. During a rate-limit store outage, both limits are counted per server.