Skip to main content
Pass a response_format and the model is constrained toward your JSON Schema during generation. Use it for anything downstream that needs a known shape: database inserts, form filling, data extraction, classification with fixed enums.

Minimal code

These examples use nemotron-3-5-lightning-30b with reasoning_effort: "none" so the model can enforce the schema without generating reasoning first.
The SDK parse() helpers validate the reply against your Pydantic or Zod model. The curl tab shows the raw response_format request underneath them, which is the portable form for any client.

Which model, and what it needs

Run GET /v1/models for the current model list and each model’s capabilities. A model that is temporarily paused is reported there and on its Model Library page.
On Nemotron 3.5 Lightning 30B, the gateway applies reasoning_effort: "none" when you omit it beside a schema. If you explicitly send an incompatible effort, the request is refused with 400, param: "response_format". Read Chat completions for model-specific behavior.

What to tune

Common mistakes

  • Asking for JSON in the prompt instead of in response_format. The model can still add commentary or markdown fences. Pass the field.
  • Leaving additionalProperties open. Without additionalProperties: false, the reply may carry keys you did not declare. The helpers set it. Raw JSON Schema users must add it.
  • Nested anyOf without discriminators. Use enums or discriminated unions. An unbounded anyOf gives the schema too little shape to enforce.
  • Expecting enums to be case-insensitive. Schema enums are exact match. "Paris" does not satisfy enum: ["paris", "berlin"].

Next steps

Tool calling

Structured arguments, then an action.

Streaming

Stream JSON that parses as it arrives.

Chat completions

The full response_format contract.