Skip to main content
The model catalog is available at https://api.pinference.ai/api/v1/models. It includes model IDs, serving types, pricing, and specifications when available. Availability and prices can change; read the current response rather than relying on a fixed list in this guide.
To select a team billing account for authenticated requests, add -H "X-Prime-Team-ID: your-team-id" (see team billing). You can also list models with prime inference models or browse the model catalog. Use the returned id as the model in a chat completion. For example, the hosted GLM-5.3 model uses z-ai/glm-5.3; see the GLM-5.3 guide.