> ## Documentation Index
> Fetch the complete documentation index at: https://runinfra.ai/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Build with RunInfra

> Call hosted open-weight models through OpenAI-compatible and Anthropic-compatible Model APIs.

RunInfra hosts open-weight models you can call from your apps and coding agents. Chat completions and Responses are OpenAI-compatible. Messages is Anthropic-compatible. Use a workspace API key and pay for usage from your credit balance.

<Note>
  Using a coding agent? Your agent can set up RunInfra's open models for you. Copy the prompt from [Set up with your coding agent](/docs/tools-sdks/agent-setup).
</Note>

## Call a hosted model

Set the base URL to `https://api.runinfra.ai/v1`, use a workspace API key, and pass a model id. There is nothing to provision.

<CodeGroup>
  ```python Python theme={"dark"}
  import os
  from openai import OpenAI

  client = OpenAI(
      base_url="https://api.runinfra.ai/v1",
      api_key=os.environ["RUNINFRA_GATEWAY_KEY"],
  )

  response = client.chat.completions.create(
      model="nemotron-3-5-lightning-30b",
      messages=[{"role": "user", "content": "Write a TypeScript retry helper."}],
      max_tokens=16384,
  )
  print(response.choices[0].message.content)
  ```

  ```javascript TypeScript theme={"dark"}
  import OpenAI from "openai";

  const client = new OpenAI({
    baseURL: "https://api.runinfra.ai/v1",
    apiKey: process.env.RUNINFRA_GATEWAY_KEY,
  });

  const response = await client.chat.completions.create({
    model: "nemotron-3-5-lightning-30b",
    messages: [{ role: "user", content: "Write a TypeScript retry helper." }],
    max_tokens: 16384,
  });
  console.log(response.choices[0].message.content);
  ```

  ```bash curl theme={"dark"}
  curl https://api.runinfra.ai/v1/chat/completions \
    -H "Authorization: Bearer $RUNINFRA_GATEWAY_KEY" \
    -H "Content-Type: application/json" \
    -d '{"model":"nemotron-3-5-lightning-30b","messages":[{"role":"user","content":"Write a TypeScript retry helper."}],"max_tokens":16384}'
  ```
</CodeGroup>

`GET /v1/models` is the authoritative list of hosted chat models. It also reports when one is temporarily paused.

<Columns cols={4}>
  <Card title="Model APIs quickstart" icon="rocket" href="/docs/api-reference/model-apis-quickstart">
    Key, base URL, first call.
  </Card>

  <Card title="OpenAI compatibility" icon="plug" href="/docs/tools-sdks/openai-compatibility">
    LangChain, LiteLLM, LlamaIndex, AI SDK, plain fetch.
  </Card>

  <Card title="Limits" icon="gauge" href="/docs/api-reference/rate-limits">
    Request, concurrency, and token limits by operation.
  </Card>

  <Card title="Coding agents" icon="terminal" href="/docs/tools-sdks/agent-setup">
    Copy one prompt. Your coding agent sets up RunInfra's open models for you.
  </Card>
</Columns>

## Money, keys, and when it breaks

<Columns cols={3}>
  <Card title="Pricing and credits" icon="credit-card" href="/docs/introduction/plans">
    Pay as you go, \$5 minimum top-up, and monthly coding plans.
  </Card>

  <Card title="Account and access" icon="key" href="/docs/faq/account">
    Sign up, keys, workspaces, deletion.
  </Card>

  <Card title="Troubleshooting" icon="life-buoy" href="/docs/tips/troubleshooting">
    401, 402, 403, 429, 503, and what to do about each.
  </Card>
</Columns>
