z.ai

Open Source

GLM-4.7-FlashXpricing and what it's actually good for

GLM-4.7-FlashX from z.ai. GLM weights are open, so third-party hosts often undercut z.ai's own first-party API — the spread is real buying information and is worth publishing alongside this rate, attributed. z.ai lists cached input storage as limited-time free.

API identifier
glm-4.7-flashx

Pricing · per 1M tokens

Input$0.07
Output$0.40
Cached input$0.01

Cached input is 7× cheaper on this model — it pays off when you send the same context again and again.

Checked against z.ai pricing docs on 7 Aug 2026, 11:51 UTC

Not sure GLM-4.7-FlashX fits your workload?

We'll size the right model for your use case and give you a real cost estimate. One free call, no slides.

Talk to us