Anthropic released Claude Opus 5.5 on September 22, 2026, the first model in a new 5.5 family. It performs at the level of Claude Fable 5.1 on most work while costing 40 percent less to run than Opus 5, and generates output more than 30 percent faster. API pricing drops to 4 dollars per million input tokens and 20 dollars per million output tokens, down from 5 and 25. It is available on all platforms, including AWS, Google Cloud, and Microsoft Azure.
What actually improved?
On benchmarks Anthropic published, Opus 5.5 surpassed Fable 5.1 on agentic coding, knowledge work, computer use, visual chart recognition, and multidisciplinary reasoning. The named benchmarks include Terminal-Bench 4.0, FrontierCode v1.1 Main, CursorBench 4.0, and GDPval-AA v2.1. Anthropic cites one early tester completing a 680,000-line code migration in under a day, work it says would have taken an engineering team weeks.
There is also a communication change that is easy to overlook. Anthropic says Opus 5.5 uses less jargon and puts important information at the start of a response rather than building to it, so answers are easier to scan. That is a behavioural change rather than a capability one, but for anyone reading model output all day it may matter more than a benchmark point.
How does the pricing actually break down?
- Input: 4 dollars per million tokens, down 20 percent from 5 dollars
- Output: 20 dollars per million tokens, down 20 percent from 25 dollars
- Cache write: 5 dollars per million tokens, down from 6.25 dollars
- Cache read: 0.20 dollars per million tokens, down sharply from 0.50 dollars
The headline 40 percent figure is not the same as the 20 percent rate cut. Anthropic arrives at 40 percent by combining the lower rates with the claim that the model uses fewer tokens to finish a given task. That is a real effect if it holds on your workload, but it is a per-task estimate rather than a posted price, so teams should measure it against their own usage rather than assume it.
Why does the safety framing matter here?
Anthropic explicitly positions this as its first release since calling for pacing the frontier, referring to CEO Dario Amodei’s essay urging the industry to slow capability advancement, which we covered in our report on the slowdown call and the safety resignations that preceded it. The company says Opus 5.5 was tested before release by external evaluators including Frontier Design and METR, and that on its automated behavioral audit, the most comprehensive alignment test it runs, Opus 5.5 is the strongest-performing model it has tested.
The obvious tension is that shipping a more capable model two months after Opus 5 is not, on its face, slowing down. Anthropic’s answer is that pacing refers to the rate of capability advancement paired with safety work, not a freeze on releases. Whether that distinction satisfies the researchers who resigned over exactly this question is a separate matter, and one worth watching as Sonnet 5.5 and Haiku 5.5 arrive.
Because Opus 5.5 is comparable to Mythos 5.1 in biology and cybersecurity capability, it ships under the same safeguards applied to Fable 5.1. Those limit how far the model can be pushed on discovering exploits in compiled programs or developing recognizable biological weapons, the same restricted-access logic behind Google’s decision to gate its most capable cybersecurity model.
What changes for subscribers?
Anthropic raised the five-hour usage limits on Pro, Max, Team, and seat-based Enterprise plans alongside the launch. Subscribers also received a rate limit reset usable any time between now and October 22. Sonnet 5.5 and Haiku 5.5 are promised in the coming weeks with many of the same improvements to performance, efficiency, and safety.
What is the honest caveat?
Every benchmark figure above comes from Anthropic’s own evaluations. Independent verification has not caught up yet, and the industry has a pattern of vendor-published numbers looking stronger than third-party results, a problem our guide to how AI model rankings work covers in detail. The external evaluator involvement from METR and Frontier Design applies to safety testing, not to the performance claims. Treat the capability numbers as a starting point for your own testing rather than a settled result.
See Anthropic’s own announcement and TechCrunch’s coverage for the full detail.




