All models

glm-5.3-flash

glm-5.3-flash
StreamingTool callingVisionReasoning
Get API key

Overview

glm-5.3-flash is available on MAX API through the OpenAI-compatible API.

Pricing

Official list prices in USD. Token prices are per 1M tokens.
TierInputOutputCache readCache write
All requests$0.15$0.5$0.03$0

Cache read applies to prompt tokens served from the prompt cache; cache write applies to tokens written into it.

Example request

curl https://clubs.byte-ai.cn/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "glm-5.3-flash",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'