glm-5.3-flash
BytePlus
glm-5.3-flash
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.
Input / output per 1M
$0.15 / $0.50
Context Window
1M
Max output
128,000
Released
Aug 26, 2026
Provider pricing
The same model is offered by several providers. Prices below are per provider; select a row to switch provider.
| Provider | Input / 1M | Output / 1M | Cache read / 1M | Context | Max output |
|---|---|---|---|---|---|
Current | $0.15 | $0.50 | $0.03 | 1M | 128,000 |
Pricing
Prices from the current provider, per 1M tokens.
| Input / 1M | Output / 1M | Cache read / 1M |
|---|---|---|
| $0.15 | $0.50 | $0.03 |
Performance
Data for this section is coming soon.
Uptime
Data for this section is coming soon.