Skip to main content
Send OpenAI-compatible requests to POST https://api.pinference.ai/api/v1/chat/completions. Pass an exact model ID from the models API.
To charge a team, add -H "X-Prime-Team-ID: your-team-id" to each request. See team billing for SDK setup and account details.

Stream a response

Set "stream": true in the request body to receive server-sent events as tokens are generated. In the OpenAI Python SDK:

Usage details

If you need token counts and cost in the response, request usage details with the Prime Inference extension:
Supported parameters and limits depend on the selected model. Check the model catalog for current capabilities rather than assuming every model accepts the same options.