DoublewordDoubleword

Model Name

zai-org/GLM-5.3-Flash

GLM-5.3-Flash

  • Type: Generation
  • Capabilities: reasoning
  • Cache read: $0.03 per 1M input tokens (0.2× standard input price). See prompt caching.

Overview

GLM-5.3-Flash is Z.ai’s natively multimodal model for coding, reasoning, and agentic workflows. It has 320 billion total parameters with 18 billion active per token, combining sparse and linear attention for efficient long-context processing. Trained on a 30-trillion-token multimodal corpus, it improves on GLM-5.2 across coding, tool-use, and general capability benchmarks while requiring substantially less serving compute.

Reasoning efforts

  • Supported: minimal, low, medium, high, xhigh, max

See the reasoning effort guide for request examples.

Pricing

PriorityInput Tokens (per 1M)Cache Read Tokens (per 1M)Output Tokens (per 1M)
Realtime1$0.15$0.03$0.50
Async$0.11$0.02$0.38
Batch (24h)$0.08$0.02$0.25

Playground

Open this model in the Playground.

Footnotes

  1. Realtime availability is limited. Doubleword is primarily a batch API.