Skip to main content
GLM-5.3 is Z.ai’s open-weights reasoning model for coding and long-running agent tasks. Prime Inference serves z-ai/glm-5.3 on Prime GPU infrastructure as a hosted model. See the model catalog or the Models API for current pricing and availability. Other GLM variants have different capabilities and serving providers.

Use Prime Inference

Create a Prime API key with Inference permission in the API keys guide, then run:
To charge a team instead of your personal account, add -H "X-Prime-Team-ID: your-team-id". See the Inference quickstart for Python SDK setup and billing details. Keep API keys server-side.