> ## Documentation Index
> Fetch the complete documentation index at: https://runinfra.ai/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Structured output

> Constrain a reply to a JSON schema with response_format, on the models that can enforce one.

Pass a `response_format` and the model is constrained toward your JSON Schema during generation. Use it for anything downstream that needs a known shape: database inserts, form filling, data extraction, classification with fixed enums.

## Minimal code

These examples use `nemotron-3-5-lightning-30b` with `reasoning_effort: "none"` so the model can enforce the schema without generating reasoning first.

<CodeGroup>
  ```python Python theme={"dark"}
  import os
  from openai import OpenAI
  from pydantic import BaseModel

  client = OpenAI(
      base_url="https://api.runinfra.ai/v1",
      api_key=os.environ["RUNINFRA_GATEWAY_KEY"],
  )

  class Receipt(BaseModel):
      merchant: str
      total_usd: float
      date: str
      items: list[str]

  response = client.beta.chat.completions.parse(
      model="nemotron-3-5-lightning-30b",
      messages=[
          {"role": "system", "content": "Extract the receipt fields."},
          {"role": "user", "content": "I bought coffee and a muffin at Blue Bottle on 2026-04-20 for $12.50"},
      ],
      reasoning_effort="none",
      response_format=Receipt,
  )

  receipt: Receipt = response.choices[0].message.parsed
  print(receipt.merchant, receipt.total_usd)
  ```

  ```typescript TypeScript theme={"dark"}
  import OpenAI from "openai";
  import { z } from "zod";
  import { zodResponseFormat } from "openai/helpers/zod";

  const client = new OpenAI({
    baseURL: "https://api.runinfra.ai/v1",
    apiKey: process.env.RUNINFRA_GATEWAY_KEY,
  });

  const Receipt = z.object({
    merchant: z.string(),
    total_usd: z.number(),
    date: z.string(),
    items: z.array(z.string()),
  });

  const response = await client.chat.completions.parse({
    model: "nemotron-3-5-lightning-30b",
    messages: [
      { role: "system", content: "Extract the receipt fields." },
      { role: "user", content: "I bought coffee and a muffin at Blue Bottle on 2026-04-20 for $12.50" },
    ],
    reasoning_effort: "none",
    response_format: zodResponseFormat(Receipt, "receipt"),
  });

  const receipt = response.choices[0].message.parsed;
  console.log(receipt?.merchant, receipt?.total_usd);
  ```

  ```bash cURL theme={"dark"}
  curl https://api.runinfra.ai/v1/chat/completions \
    -H "Authorization: Bearer $RUNINFRA_GATEWAY_KEY" \
    -H "Content-Type: application/json" \
    -d '{
      "model": "nemotron-3-5-lightning-30b",
      "messages": [
        {"role":"system","content":"Extract the receipt fields."},
        {"role":"user","content":"I bought coffee at Blue Bottle on 2026-04-20 for $12.50"}
      ],
      "reasoning_effort": "none",
      "response_format": {
        "type": "json_schema",
        "json_schema": {
          "name": "receipt",
          "strict": true,
          "schema": {
            "type": "object",
            "properties": {
              "merchant":  {"type":"string"},
              "total_usd": {"type":"number"},
              "date":      {"type":"string"},
              "items":     {"type":"array","items":{"type":"string"}}
            },
            "required": ["merchant","total_usd","date","items"],
            "additionalProperties": false
          }
        }
      }
    }'
  ```
</CodeGroup>

The SDK `parse()` helpers validate the reply against your Pydantic or Zod model. The curl tab shows the raw `response_format` request underneath them, which is the portable form for any client.

## Which model, and what it needs

<Note>
  Run `GET /v1/models` for the current model list and each model's capabilities. A model that is temporarily paused is reported there and on its [Model Library](https://runinfra.ai/inference-api) page.
</Note>

On Nemotron 3.5 Lightning 30B, the gateway applies `reasoning_effort: "none"` when you omit it beside a schema. If you explicitly send an incompatible effort, the request is refused with `400`, `param: "response_format"`. Read [Chat completions](/docs/api-reference/chat-completions#three-things-a-request-can-be-refused-for) for model-specific behavior.

## What to tune

| Parameter | Effect |
| - | - |
| `response_format: { type: "json_object" }` | Looser. Any valid JSON, no schema enforcement. |
| `response_format.json_schema.strict: true` | Accepted for OpenAI compatibility. The Pydantic and Zod helpers set it for you. |
| `temperature: 0` | Best for deterministic extraction. Non-zero works too. |
| `reasoning_effort: "none"` | Lets Nemotron 3.5 Lightning 30B enforce a schema without reasoning tokens. The gateway applies it when you omit the field beside a schema. |

## Common mistakes

* **Asking for JSON in the prompt instead of in `response_format`.** The model can still add commentary or markdown fences. Pass the field.
* **Leaving `additionalProperties` open.** Without `additionalProperties: false`, the reply may carry keys you did not declare. The helpers set it. Raw JSON Schema users must add it.
* **Nested `anyOf` without discriminators.** Use enums or discriminated unions. An unbounded `anyOf` gives the schema too little shape to enforce.
* **Expecting enums to be case-insensitive.** Schema enums are exact match. `"Paris"` does not satisfy `enum: ["paris", "berlin"]`.

## Next steps

<Columns cols={3}>
  <Card title="Tool calling" icon="wrench" href="/docs/cookbook/tool-calling">
    Structured arguments, then an action.
  </Card>

  <Card title="Streaming" icon="zap" href="/docs/cookbook/streaming">
    Stream JSON that parses as it arrives.
  </Card>

  <Card title="Chat completions" icon="square-terminal" href="/docs/api-reference/chat-completions">
    The full `response_format` contract.
  </Card>
</Columns>
