Pricing and documentation verified: 20 August 2026. Compiled from Anthropic’s official published documentation — the API pricing page, the models overview comparison table, the token counting guide and the context windows guide — all read on 20 August 2026. We do not hold an Anthropic API account and have not run any of these models or benchmarked them ourselves. Every price and every word-and-character figure below is a published vendor number. Every ratio, per-word rate and worked example built on top of them is our own arithmetic and is labelled where it appears.
Compiled and fact-checked by the AI Tools Worth editorial team. Corrections: contact page.
The sticker price stopped being the price
Anthropic’s per-token rates are published, stable and easy to compare. Claude Sonnet 5 costs $2 per million input tokens. Claude Sonnet 4.6 cost $3. That reads as a 33.3% price cut, and every comparison table on the internet reports it that way.
It is not a 33.3% cut for the same text, because the two models do not count text the same way. Anthropic states this plainly in three separate places in its own documentation. From the pricing page:
“Claude 4.7 and later models and Claude Mythos Preview use a newer tokenizer that contributes to their improved performance on a wide range of tasks. This tokenizer produces approximately 30% more tokens for the same text. The exact increase depends on the content and workload shape. Claude Sonnet 4.6 and earlier models use the previous tokenizer.”
A per-token price is only comparable across models that agree on what a token is. Once the tokenizer changes underneath, the honest unit of comparison is cost per unit of text, not cost per token. That is what this page works out, using Anthropic’s own published word and character equivalences.
Anthropic’s own numbers put the inflation at 35–36%, not 30%
The prose says “approximately 30%”. The model comparison table quantifies it more precisely, in the tooltips attached to each model’s context window figure. Those tooltips give a word count and a character count for the same 1M-token window, and they split cleanly into two groups.
| Tokenizer | Models | 1M tokens holds | Implied ratio |
|---|---|---|---|
| Previous | Sonnet 4.6, Opus 4.6, Sonnet 4.5, Opus 4.5, Haiku 4.5 and earlier | ~750k words ~3.4M characters | 0.750 words per token 3.40 characters per token |
| Introduced with Opus 4.7 | Opus 4.7, Opus 4.8, Opus 5, Sonnet 5, Fable 5, Mythos 5 | ~555k words ~2.5M characters | 0.555 words per token 2.50 characters per token |
Run those against each other — our arithmetic — and the same text produces 35.1% more tokens measured by words (0.750 ÷ 0.555) and 36.0% more measured by characters (3.40 ÷ 2.50). Both are meaningfully above the “approximately 30%” the prose gives. We could not reconcile the two figures; see What we could not verify below. The rest of this page uses the tooltip figures, because they are the only numbers Anthropic publishes that are precise enough to compute with.
The same fact, two directions: +35% tokens, −26% capacity
This is the part that trips people up, so it is worth separating carefully. Token inflation and context-window capacity loss are the same fact stated two ways, and they are different numbers.
- Sending the same text costs 35.1% more tokens. A document that encoded to 1,000,000 tokens on Sonnet 4.6 encodes to about 1,351,000 tokens on Sonnet 5.
- The same 1M-token window holds 26.0% less text. 750k words became 555k words — a loss of 195,000 words, or roughly 26.5% by characters.
Both models are documented as having “a 1M-token context window”, and that statement is true of both. The window did not shrink. What fits inside it did, by about a quarter.
For anyone who sized a pipeline against the older figure, that is the number that matters: a chunking strategy, a retrieval budget or a document-batching limit tuned to fit 750k words into one Claude 4.6-generation request will overflow on a Claude 5-generation model, and it will overflow by roughly a third more tokens than expected.
Cost per million words: the full table
This is our arithmetic: each model’s published per-million-token rate divided by the words that Anthropic says fit in a million tokens for that model’s tokenizer. It is the closest thing to an apples-to-apples price that the published figures allow.
| Model | Sticker, per MTok in / out | Words per MTok | Per M words, input | Per M words, output |
|---|---|---|---|---|
| Claude Haiku 4.5 | $1 / $5 | 750k | $1.33 | $6.67 |
| Claude Sonnet 4.5 | $3 / $15 | 750k | $4.00 | $20.00 |
| Claude Sonnet 4.6 | $3 / $15 | 750k | $4.00 | $20.00 |
| Claude Sonnet 5 | $2 / $10 | 555k | $3.60 | $18.02 |
| Claude Opus 4.5 | $5 / $25 | 750k | $6.67 | $33.33 |
| Claude Opus 4.6 | $5 / $25 | 750k | $6.67 | $33.33 |
| Claude Opus 4.7 | $5 / $25 | 555k | $9.01 | $45.05 |
| Claude Opus 4.8 | $5 / $25 | 555k | $9.01 | $45.05 |
| Claude Opus 5 | $5 / $25 | 555k | $9.01 | $45.05 |
| Claude Fable 5 | $10 / $50 | 555k | $18.02 | $90.09 |
Three readings change once the table is in per-word terms.
Sonnet 5’s 33% cut is a 9.9% cut
Sonnet 4.6 at $3 per MTok worked out to $4.00 per million words of input. Sonnet 5 at $2 per MTok works out to $3.60. That is a real cut, and it survives — but it is 9.9%, not 33.3%. The output side moves identically: $20.00 down to $18.02, again 9.9%.
Worth stating clearly, because this is not a debunking: Sonnet 5 did get cheaper per unit of text. Anthropic also confirmed on the same pricing page that the $2/$10 rate, originally announced as introductory pricing through 31 August 2026, is now the standard price and the scheduled increase to $3/$15 on 1 September will not happen — which we covered when it was announced. The cut is genuine. It is roughly a third the size the sticker suggests.
Opus 5 costs 35% more per word than Opus 4.6 at an unchanged sticker price
Opus 4.5, 4.6, 4.7, 4.8 and Opus 5 all list $5 / $25 per million tokens. Five consecutive releases, one unchanged headline number, and it is easy to read that as five generations of flat pricing.
Per million words it is not flat. Opus 4.6 was $6.67 in / $33.33 out. Opus 4.7 — the release where the tokenizer changed — is $9.01 in / $45.05 out, and Opus 4.8 and Opus 5 inherit it. That is a 35.1% increase in the cost of moving the same text, applied at a moment when the published price did not move at all.
The Haiku-to-Sonnet gap is wider than it looks
Haiku 4.5 at $1 against Sonnet 5 at $2 reads as a clean 2× step up. Because Haiku 4.5 uses the previous tokenizer and Sonnet 5 does not, the real step is 2.70× per million words ($1.33 against $3.60). Anyone routing cheap traffic to Haiku on a 2× assumption is understating the saving by about 35%.
The same distortion runs the other way across the generation boundary. Sonnet 5 at $2 against Opus 4.6 at $5 looks like a 60% saving; per word it is 45.9%.
Comparisons within a tokenizer generation are unaffected. Fable 5 is exactly 2× Opus 5 on both sticker and per-word terms, because both use the same tokenizer. The distortion only appears when a comparison straddles Opus 4.7.
Worked example: one 100,000-word document
A 100,000-word document — a long technical manual, or roughly a 400-page book — sent once as input. Token counts and costs are our arithmetic on the published rates and word ratios.
| Model | Tokenizer | Input tokens | Cost to send once |
|---|---|---|---|
| Claude Haiku 4.5 | previous | ~133,300 | $0.13 |
| Claude Sonnet 4.6 | previous | ~133,300 | $0.40 |
| Claude Sonnet 5 | new | ~180,200 | $0.36 |
| Claude Opus 4.6 | previous | ~133,300 | $0.67 |
| Claude Opus 5 | new | ~180,200 | $0.90 |
| Claude Fable 5 | new | ~180,200 | $1.80 |
The token column is the whole story. The same file is about 46,900 tokens larger on any Claude 5-generation model than it was on a 4.6-generation model, before anything else in the request is counted.
Because prompt caching and the Batch API are published as multipliers on the base rate — 1.25× and 2× for 5-minute and 1-hour cache writes, 0.1× for cache reads, 50% off for batch — every one of those derived rates shifts by the same proportion. There is no line item where the per-word effect washes out.
How to check this against your own workload
You do not have to take the 30%, or our 35.1%, on trust. Anthropic’s token counting endpoint returns the count under the tokenizer of whichever model you name, it is free to use, and the documentation explicitly describes counting the same request twice to measure the difference:
“The token counting endpoint returns the count under the tokenizer of the model you pass, so to measure the difference for your workload, count the same request twice: once with your current model and once with model: "claude-fable-5" (or "claude-mythos-5"), and compare the two input_tokens values.”
The endpoint is POST /v1/messages/count_tokens. Per the documentation it is free, carries its own rate limit separate from message creation — 2,000 requests per minute on the Start tier, 4,000 on Build, 8,000 on Scale — and accepts the same request shape as the Messages API, including system prompts, tools, images and PDFs. Send your real prompt with model: "claude-sonnet-4-6", send it again with model: "claude-sonnet-5", and the ratio between the two input_tokens values is your workload’s actual inflation figure.
That number is the one to plan against. Anthropic’s documentation twice says the increase “depends on the content and workload shape”, and published third-party measurements have reported figures ranging from roughly 12% to 35% depending on what is being encoded. We have not run this measurement ourselves and are not reporting a figure of our own.
The documentation gap worth knowing about
One detail is easy to miss. In the model comparison table, the tooltip on Claude Fable 5 explains itself in full: it states that Fable 5 uses the tokenizer introduced with Opus 4.7 and that the same text produces roughly 30% more tokens. The tooltip on Claude Opus 4.7 at least notes that it “uses a new tokenizer”.
The tooltips on Claude Opus 5, Claude Sonnet 5 and Claude Opus 4.8 carry the reduced ~555k words / ~2.5M characters figure with no tokenizer note attached. A reader comparing the Sonnet 5 and Sonnet 4.6 rows sees the word capacity drop from 750k to 555k, in a table where both rows say “1M tokens”, with nothing on either row explaining why. The explanation exists — on the pricing page, on the token counting page, and in a different model’s tooltip — but not where the discrepancy is visible.
What this does not tell you
This page measures the price of moving text, and nothing else. That is a real and comparable quantity, but it is not the same as the price of getting a job done, and the difference matters:
- A newer model may need fewer output tokens, fewer retries or fewer turns to complete the same task. Every model in the Claude 5 generation uses adaptive thinking, where the model allocates its own reasoning budget per request, so output volume varies in ways no static table can capture.
- Thinking tokens are billed as output. On Opus 4.5 and later Opus models and Sonnet 4.6 and later Sonnet models, previous thinking blocks are kept in context by default and are then billed as input on subsequent turns — a per-conversation cost that depends entirely on conversation shape.
- Per-word cost says nothing about output quality. If a more expensive model gets an answer right the first time, its higher per-word rate can still be the cheaper option.
The right way to use the table above is as a correction to sticker-price comparisons, not as a replacement for measuring your own workload.
Frequently asked questions
Did Claude Sonnet 5 actually get cheaper than Sonnet 4.6?
Yes, but by less than the headline. Per token, $2 against $3 is a 33.3% cut. Per million words of text — our arithmetic using Anthropic’s published word equivalences of 555k and 750k words per million tokens — it is $3.60 against $4.00, a 9.9% cut. The output side moves by the same 9.9%.
Which Claude models use the new tokenizer?
Claude Opus 4.7, Opus 4.8, Opus 5, Sonnet 5, Fable 5, Mythos 5 and Mythos Preview. Claude Sonnet 4.6, Opus 4.6, Sonnet 4.5, Opus 4.5, Haiku 4.5 and everything earlier use the previous tokenizer. The dividing line is Opus 4.7.
How much more does the new tokenizer cost me?
Anthropic’s prose says approximately 30% more tokens for the same text. Its own published word and character equivalences imply 35.1% and 36.0% respectively (our arithmetic). The documentation states twice that the exact figure depends on content and workload shape, so measure yours with the free token counting endpoint rather than assuming any single number.
Did Claude’s context window get smaller?
No. It is still 1M tokens on Opus 4.6 and later, Sonnet 4.6 and later, and Fable 5. What changed is how much text fits inside it: Anthropic’s tooltips put the previous-tokenizer capacity at ~750k words / ~3.4M characters and the new-tokenizer capacity at ~555k words / ~2.5M characters — about 26% less text in the same nominal window.
Why is the capacity loss 26% when token inflation is 35%?
They are the same fact inverted. If the same text needs 1.351× as many tokens, then a fixed token budget holds 1 ÷ 1.351 = 0.740 of the text it used to, which is a 26.0% reduction. Quoting 35% as a capacity loss, or 26% as a cost increase, gets the direction right and the number wrong.
Does the tokenizer change affect prompt caching and batch pricing?
Yes, proportionally. Cache writes (1.25× for 5 minutes, 2× for 1 hour), cache reads (0.1×) and the Batch API discount (50%) are all published as multipliers on the base per-token rate, so each derived rate carries the same per-word shift as the base rate does.
Does comparing Opus 5 against Fable 5 have the same problem?
No. Both use the tokenizer introduced with Opus 4.7, so Fable 5 is 2× Opus 5 on both sticker and per-word terms. The distortion only appears in comparisons that cross the Opus 4.7 boundary — for example Sonnet 4.6 against Sonnet 5, or Haiku 4.5 against anything in the Claude 5 generation.
How do I measure the inflation on my own prompts?
Send the identical request to POST /v1/messages/count_tokens twice, changing only the model field — once with a pre-4.7 model and once with your target model — and divide the two input_tokens values. The endpoint is documented as free, with rate limits of 2,000 to 8,000 requests per minute depending on usage tier, tracked separately from message creation.
What we could not verify
Gaps we could not close, listed here rather than papered over.
- Why the prose says ~30% and the tooltips imply 35–36%. Both figures are Anthropic’s, on pages linked to each other. The word and character counts are given as approximations with a leading tilde, so some of the gap is rounding, but we could find no statement reconciling them and no explanation of which is authoritative.
- What text the word and character equivalences describe. Presumably general English prose. Code, JSON, tables and non-Latin scripts tokenize differently, and the documentation does not say what corpus the ~750k and ~555k figures are drawn from.
- Whether output text inflates by the same ratio as input text. The documentation describes the tokenizer’s effect on input and on context capacity. We assume the same encoding applies to generated text, since it is the same tokenizer, but Anthropic does not state an output-side figure and we did not measure one.
- The real-world figure for any specific workload. We hold no Anthropic API account and ran no token counts. The 12–35% range mentioned above comes from third-party published measurements we did not reproduce.
- Whether the pre-4.7 word equivalence was ever revised. We read the current documentation on one date and have no archived captures of the model comparison tooltips, so we cannot show when the ~555k figure first appeared.
- Effective per-task cost. Nothing here accounts for how many tokens a given model needs to finish a given job, which is the number that actually lands on a bill.
- Partner-platform pricing. Amazon Bedrock and Google Cloud set their own rates and carry a 10% premium on regional and multi-region endpoints. We used first-party Claude API rates only.
Sources and method
Compiled 20 August 2026 from Anthropic’s official published documentation: the API pricing page, the models overview and its latest-models and legacy comparison tables, the token counting guide, and the context windows guide. Per-token rates, the tokenizer note, the word and character equivalences, the cache and batch multipliers, and the token counting endpoint’s rate limits and free-of-charge status are all quoted or taken directly from those pages as they stood on that date. All per-word rates, ratios, percentage changes and worked examples are our own arithmetic on those published figures and are labelled where they appear; because the word and character equivalences are published as approximations, the arithmetic inherits their rounding. We hold no Anthropic API account, have not run these models, and have not measured token counts ourselves. Prices and documentation change — verify against the vendor before you commit.
Assessed verdicts, real price changes and the launches that matter. No hype, no spam — unsubscribe anytime.