> ## Documentation Index
> Fetch the complete documentation index at: https://runinfra.ai/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Hugging Face clients

> Call RunInfra Model APIs from huggingface_hub and @huggingface/inference by passing a base URL, with the verified model id mapping and capability matrix.

Both Hugging Face inference clients, `huggingface_hub` for Python and `@huggingface/inference` for JavaScript, accept an OpenAI-compatible base URL, so both call RunInfra today.

RunInfra is not a registered Hugging Face Inference Provider, so `provider="runinfra"` does not resolve. Use the base-URL forms below.

## Python

```bash theme={"dark"}
pip install "huggingface_hub>=1.0"
export RUNINFRA_GATEWAY_KEY="rp_your_key_here"
```

```python theme={"dark"}
import os
from huggingface_hub import InferenceClient

client = InferenceClient(
    base_url="https://api.runinfra.ai/v1",
    api_key=os.environ["RUNINFRA_GATEWAY_KEY"],
)

completion = client.chat.completions.create(
    model="nemotron-3-5-lightning-30b",
    messages=[{"role": "user", "content": "Write a TypeScript retry helper."}],
    max_tokens=16384,
)

print(completion.choices[0].message.content)
```

Streaming is the same call with `stream=True`, reading `chunk.choices[0].delta.content` and skipping the usage frame, which has no choices.

## JavaScript

Pass `endpointUrl` in the client options and your RunInfra key as the token. The client appends `/v1/chat/completions` for you, so give it the host with or without `/v1`.

```bash theme={"dark"}
npm install @huggingface/inference
export RUNINFRA_GATEWAY_KEY="rp_your_key_here"
```

```javascript theme={"dark"}
import { InferenceClient } from "@huggingface/inference";

const client = new InferenceClient(process.env.RUNINFRA_GATEWAY_KEY, {
  endpointUrl: "https://api.runinfra.ai",
});

const completion = await client.chatCompletion({
  model: "nemotron-3-5-lightning-30b",
  messages: [{ role: "user", content: "Write a TypeScript retry helper." }],
  max_tokens: 16384,
});

console.log(completion.choices[0].message.content);
```

`chatCompletionStream` takes the same arguments and yields deltas.

<Warning>
  Do not combine `endpointUrl` with an explicit third-party `provider`. The client rejects that combination with `Cannot use endpointUrl with a third-party provider`. Pass `endpointUrl` alone.
</Warning>

## Repository mappings

These clients normally take a Hugging Face repository id. In the base-URL forms above you pass the RunInfra model id, because you are talking to RunInfra directly.

These mappings were available on September 21, 2026. Run `GET /v1/models` for the current list; a model that is temporarily paused is reported there and on its Model Library page.

| RunInfra model id | Hugging Face repository |
| - | - |
| `nemotron-3-5-lightning-30b` | `nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16` |
| `ornith-1-5-35b` | `ornith-ai/Ornith-1.5-35B-A3B` |
| `qwen3-8-27b` | `Qwen/Qwen3.8-27B` |

For the mappings above, the gateway also accepts the repository id in the `model` field, so either string resolves. `GET /v1/models` returns the RunInfra ids and is the authoritative live list.

## What each model does

<Note>
  Run `GET /v1/models` for the current model list and each model's capabilities. A model that is temporarily paused is reported there and on its [Model Library](https://runinfra.ai/inference-api) page.
</Note>

Working JSON snippets are in [Structured output](/docs/cookbook/structured-output). When a model needs a specific `reasoning_effort` for JSON output, the gateway applies it if you omit the field.

## Where the model's thinking arrives

These models think before they answer. That text streams on `delta.reasoning` and appears on `message.reasoning`, never on `content`.

## Which client should you use

If you already live inside the Hugging Face ecosystem, the clients on this page work and nothing here changes later. Otherwise use the OpenAI client, which is the shortest path and what the rest of these docs assume. See [OpenAI compatibility](/docs/tools-sdks/openai-compatibility).

Either way, calls are billed by RunInfra against your workspace balance at the per-token price on each model's page. See [Pricing and credits](/docs/introduction/plans) and [Rate limits](/docs/api-reference/rate-limits).

## Related

<Columns cols={3}>
  <Card title="OpenAI compatibility" icon="plug" href="/docs/tools-sdks/openai-compatibility">
    The shortest path, and what the rest of these docs assume.
  </Card>

  <Card title="Chat completions" icon="braces" href="/docs/api-reference/chat-completions">
    The full request contract.
  </Card>

  <Card title="Pricing and credits" icon="credit-card" href="/docs/introduction/plans">
    How these calls are billed.
  </Card>
</Columns>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.