GLM-4.5-Flash
GLM-4.5-Flash: a glm-flash model from Zhipu AI, ~131.1K context, knowledge cutoff 2025-04
GLM-4.5-Flash is a glm-flash model from Zhipu AI (~131.1K context, input around $0/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation
starsCapabilities
codeFunction calling
paymentsContext and pricing
Context limit131,072
Max output98,304
Knowledge cutoff2025-04
Input price$0/ 1M tokens
Output price$0/ 1M tokens
Cached input price$0/ 1M tokens
descriptionOverview
Overview
GLM-4.5-Flash is provided by Zhipu AI, model ID glm-4.5-flash
Key specs
- Context: 131.1K tokens
- Max output: 98.3K tokens
- Knowledge cutoff: 2025-04
- Input price: $0/1M
- Output price: $0/1M
Best for
Consider GLM-4.5-Flash when comparing context length, pricing, multimodal support and relay availability
lightbulbUse cases
- Assistants and customer support
- Content generation and rewriting
- Knowledge Q&A and summarization
- Structured extraction
thumb_upStrengths
- Large context window (~131.1K)
- Open weights, can be self-hosted
- High single-response output limit (~98.3K)
infoLimitations
- Knowledge cutoff 2025-04; newer facts need external retrieval
- Pricing and availability vary by upstream and relay; verify with official docs and tests
Scan to join WeChat group