Model Name
zai-org/GLM-5.3GLM 5.3
- Type: Generation
- Capabilities:
reasoning,enhanced_structured_generation - Cache read: $0.28 per 1M input tokens (0.2× standard input price). See prompt caching.
Overview
GLM 5.3 is Z.ai's open-weight reasoning model for complex coding, long-horizon agentic work, tool use, cybersecurity analysis and repository-scale engineering. It provides a 1M-token context window and configurable reasoning effort.
Pricing
| Priority | Input Tokens (per 1M) | Cache Read Tokens (per 1M) | Output Tokens (per 1M) |
|---|---|---|---|
| Realtime1 | $1.40 | $0.28 | $4.40 |
| Async | $1.05 | $0.21 | $3.30 |
| Batch (24h) | $0.70 | $0.14 | $2.20 |
Playground
Open this model in the Playground.
Footnotes
-
Realtime availability is limited. Doubleword is primarily a batch API. ↩