Skip to main content
Both Hugging Face inference clients, huggingface_hub for Python and @huggingface/inference for JavaScript, accept an OpenAI-compatible base URL, so both call RunInfra today. RunInfra is not a registered Hugging Face Inference Provider, so provider="runinfra" does not resolve. Use the base-URL forms below.

Python

Streaming is the same call with stream=True, reading chunk.choices[0].delta.content and skipping the usage frame, which has no choices.

JavaScript

Pass endpointUrl in the client options and your RunInfra key as the token. The client appends /v1/chat/completions for you, so give it the host with or without /v1.
chatCompletionStream takes the same arguments and yields deltas.
Do not combine endpointUrl with an explicit third-party provider. The client rejects that combination with Cannot use endpointUrl with a third-party provider. Pass endpointUrl alone.

Repository mappings

These clients normally take a Hugging Face repository id. In the base-URL forms above you pass the RunInfra model id, because you are talking to RunInfra directly. These mappings were available on September 21, 2026. Run GET /v1/models for the current list; a model that is temporarily paused is reported there and on its Model Library page. For the mappings above, the gateway also accepts the repository id in the model field, so either string resolves. GET /v1/models returns the RunInfra ids and is the authoritative live list.

What each model does

Run GET /v1/models for the current model list and each model’s capabilities. A model that is temporarily paused is reported there and on its Model Library page.
Working JSON snippets are in Structured output. When a model needs a specific reasoning_effort for JSON output, the gateway applies it if you omit the field.

Where the model’s thinking arrives

These models think before they answer. That text streams on delta.reasoning and appears on message.reasoning, never on content.

Which client should you use

If you already live inside the Hugging Face ecosystem, the clients on this page work and nothing here changes later. Otherwise use the OpenAI client, which is the shortest path and what the rest of these docs assume. See OpenAI compatibility. Either way, calls are billed by RunInfra against your workspace balance at the per-token price on each model’s page. See Pricing and credits and Rate limits.

OpenAI compatibility

The shortest path, and what the rest of these docs assume.

Chat completions

The full request contract.

Pricing and credits

How these calls are billed.