z.ai
Open Source
GLM-5-Turbo — pricing and what it's actually good for
GLM-5-Turbo from z.ai. GLM weights are open, so third-party hosts often undercut z.ai's own first-party API — the spread is real buying information and is worth publishing alongside this rate, attributed. z.ai lists cached input storage as limited-time free.
- Released
- 15 Mar 2026
- API identifier
- glm-5-turbo
Pricing · per 1M tokens
Input$1.20
Output$4.00
Cached input$0.24
Cached input is 5× cheaper on this model — it pays off when you send the same context again and again.
Checked against z.ai pricing docs on 7 Aug 2026, 11:51 UTC
Not sure GLM-5-Turbo fits your workload?
We'll size the right model for your use case and give you a real cost estimate. One free call, no slides.
Talk to us