z-ai/glm-5.3 on Prime GPU infrastructure as a hosted model.
See the model catalog or the Models API for current pricing and availability. Other GLM variants have different capabilities and serving providers.
Use Prime Inference
Create a Prime API key with Inference permission in the API keys guide, then run:-H "X-Prime-Team-ID: your-team-id". See the Inference quickstart for Python SDK setup and billing details. Keep API keys server-side.