z.ai
Open Source
GLM-5 — pricing, context window and what it's actually good for
GLM-5 from z.ai. GLM weights are open, so third-party hosts often undercut z.ai's own first-party API — the spread is real buying information and is worth publishing alongside this rate, attributed. z.ai lists cached input storage as limited-time free.
- Context window
- 200,000 tokens
- Max output
- 128,000 tokens
- Released
- 12 Feb 2026
- API identifier
- glm-5
Pricing · per 1M tokens
Input$1.00
Output$3.20
Cached input$0.20
Cached input is 5× cheaper on this model — it pays off when you send the same context again and again.
Checked against z.ai pricing docs on 7 Aug 2026, 11:51 UTC
Used for
Where GLM-5 shows up in practice
The workflows where we shortlist GLM-5 against the alternatives. Each page says why, at what price, and where a cheaper model does the job instead.
Not sure GLM-5 fits your workload?
We'll size the right model for your use case and give you a real cost estimate. One free call, no slides.
Talk to us