Proprietary
Claude Sonnet 4.6 — pricing, context window and what it's actually good for
Claude Sonnet 4.6 is an Active Claude model on the Claude API. 1M-token context, 128k max output. Training data cutoff January 2026. Uses the tokenizer used by Claude Sonnet 4.6 and earlier, so per-token costs are directly comparable with models of that generation. Tentative retirement date: not sooner than 17 February 2027 (Anthropic model deprecations, checked 2026-08-07 11:51 UTC).
- Context window
- 1,000,000 tokens
- Max output
- 128,000 tokens
- Released
- 17 Feb 2026
- Knowledge cutoff
- Aug 2025
- API identifier
- claude-sonnet-4-6
Pricing · per 1M tokens
Cached input is 10× cheaper on this model — it pays off when you send the same context again and again.
US data residency · 1.1× list
Pinning inference to the US multiplies every token category — input, output, cache writes and cache reads — by 1.1. Global routing is the default and bills at the rates above. Anthropic pricing docs add that the partner-operated platforms — Amazon Bedrock and Google Cloud — have independent regional pricing, so check theirs rather than assuming these figures carry over.
Checked against Anthropic pricing docs on 7 Aug 2026, 11:51 UTC
Used for
Where Claude Sonnet 4.6 shows up in practice
The workflows where we shortlist Claude Sonnet 4.6 against the alternatives. Each page says why, at what price, and where a cheaper model does the job instead.
Not sure Claude Sonnet 4.6 fits your workload?
We'll size the right model for your use case and give you a real cost estimate. One free call, no slides.
Talk to us