Skip to content

glm-5.3-flash

BytePlus

glm-5.3-flash

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.

Input / output per 1M
$0.15 / $0.50
Context Window
1M
Max output
128,000
Released
Aug 26, 2026
Provider pricing

The same model is offered by several providers. Prices below are per provider; select a row to switch provider.

ProviderInput / 1MOutput / 1MCache read / 1MContextMax output
Current
$0.15$0.50$0.03

1M

128,000

Pricing

Prices from the current provider, per 1M tokens.

Input / 1MOutput / 1MCache read / 1M
$0.15$0.50$0.03
Performance

Data for this section is coming soon.

Uptime

Data for this section is coming soon.

One API for the world’s leading AI models.

© 2026 Cloudsome. All rights reserved.