What changed
Anthropic released Claude Opus 5.5 on September 22, 2026: on the Claude Platform as claude-opus-5-5, and on Amazon Web Services, Google Cloud and Microsoft Azure 1. Input and output tokens drop 20% below Opus 5, and cache reads drop 60% 1. The headline saving is scoped: at default settings, 40% less than Opus 5 on typical workloads 1.
Where the discount lands
| Price per 1M tokens | Opus 5 | Opus 5.5 |
|---|---|---|
| Cache reads 1 | $0.50 | $0.20 |
| Input 1 | $5 | $4 |
| Output 1 | $25 | $20 |
The uneven line is the one that matters. Anthropic says cache reads make up the majority of agentic and coding work costs 1, and Simon Willison notes that in longer agentic conversations 90%+ of input tokens are processed at cached prices 3. Artificial Analysis puts the new cache-read price at a 95% discount on uncached input, up from 90% on previous Opus models 2. I think the second-order effect matters more: a cache miss now costs relatively more, so anything that breaks your prefix (a timestamp in the system prompt, a reordered tool list) eats a bigger share of the saving.
Max effort spends the discount
The case for max is real. At max effort, Artificial Analysis scores Opus 5.5 at 58 on its Intelligence Index, the highest score it has measured by several points 2. The bill follows. At max, Opus 5.5 uses about 119k output tokens per Intelligence Index task against about 73k for Opus 5 at max, which leaves it level with Opus 5 on cost per task 2. The 40% belongs to default settings, where the effort is medium 1.
Simon Willison ran max on one prompt and hit the 128,000 maximum output token limit (shared by the other Claude models) while the model was still reasoning, twice 3. Each failure cost him $2.56 and nearly 20 minutes 3. He suspects max is effectively useless 3. On one prompt I would not go that far, but I would not make it a default either.
Impact on your team
Opus 5.5 is no longer available with thinking mode switched off 1, so any request that switches thinking off has to change before you move to claude-opus-5-5. If your API account was created on or after August 31, 2026, preserved thinking applies; it stops API users from editing Claude’s prior context in an attempt to extract Claude’s reasoning 1. I would test any harness that rewrites or trims earlier assistant turns first. Budget on medium, the default, and route a task to max only where your evals show the gain: at max, Artificial Analysis measures no per-task saving against Opus 5 at max 2. And track your cache hit rate: your share of the cut depends on it.