z.ai

Open Source

GLM-5-Turbopricing and what it's actually good for

GLM-5-Turbo from z.ai. GLM weights are open, so third-party hosts often undercut z.ai's own first-party API — the spread is real buying information and is worth publishing alongside this rate, attributed. z.ai lists cached input storage as limited-time free.

Released
15 Mar 2026
API identifier
glm-5-turbo

Pricing · per 1M tokens

Input$1.20
Output$4.00
Cached input$0.24

Cached input is 5× cheaper on this model — it pays off when you send the same context again and again.

Checked against z.ai pricing docs on 7 Aug 2026, 11:51 UTC

Not sure GLM-5-Turbo fits your workload?

We'll size the right model for your use case and give you a real cost estimate. One free call, no slides.

Talk to us