Anthropic launches Claude Opus 5.5: Fable 5.1 performance at 40% lower cost
AI-generated
1. Executive Summary
Anthropic has unveiled Claude Opus 5.5, the first model in its new Claude 5.5 family. The company states that its performance is on a par with Claude Fable 5.1 for most tasks and estimates that it offers a 40 per cent reduction in running costs compared with Opus 5. This is Anthropic’s first release since its public call to ‘set the pace at the frontier’, a stance through which the company is calling for a more measured approach to the roll-out of the most capable models.
The model was evaluated prior to publication by external assessors, including Frontier Design and METR. In Anthropic’s automated behavioural audit — the most comprehensive alignment test the company carries out — Opus 5.5 achieved the highest score recorded to date by an Anthropic model. The company is deploying it with the safeguards reserved for its most capable models.
The launch includes three stated areas of improvement: performance, security and cost-per-speed ratio. Anthropic has also confirmed that Claude Sonnet 5.5 and Claude Haiku 5.5 will be released in the coming weeks with equivalent improvements in efficiency and security.
2. Technical Analysis: Performance and Benchmarks
Anthropic has published a series of benchmarks in which Opus 5.5 leads the categories of agentic coding, computer use and knowledge work. The company points out, however, that at these levels of capability, the margins between benchmark scores are no longer a reliable indicator of actual differences: in its own internal use, the gap between Opus 5.5 and Claude Fable 5.1 is smaller than the scores suggest.
| Area and benchmark | Opus 5.5 | Fable 5.1 | Opus 5 | GPT-6 Astra |
|---|---|---|---|---|
| Agent-based modelling · Terminal-Bench 4.0 | 66.4 per cent | 55.8 per cent | 52.3 per cent | 57.9 per cent |
| Agent-based coding · FrontierCode v1.1 | 54.4 per cent | 50.3 per cent | 48.0 per cent | 53.3 per cent |
| Agent-based modelling · CursorBench 4.0 | 57.8 per cent | 51.8 per cent | 46.6 per cent | — |
| Knowledge base · GDPval-AA v2.1 | 1846 | 1735 | 1708 | 1542 |
| Business Workflows · AutomationBench | 40.0 per cent | 31.4 per cent | 26.9 per cent | 41.4 per cent |
| Multidisciplinary reasoning · Humanity’s Last Exam | 67.7 per cent | 65.6 per cent | 63.6 per cent | 57.2 per cent |
| Agent-based scientific research · Terminal-Bench-Science 0.1 | 58.7 per cent | 52.6 per cent | 29.0 per cent | 64.6 per cent |
| Computer use · OSWorld 2.0 | 81.8 per cent | 80.7 per cent | 74.0 per cent | — |
| Visual recognition of charts · Chartography | 89.0 per cent | 88.4 per cent | 83.4 per cent | — |
Anthropic backs up the figures with use cases reported by early testers. One of them completed a code migration of 680,000 lines in less than a day – a task the company estimates would have taken an engineering team weeks to complete. In a test to optimise load times across all pages of a web application, Opus 5.5 succeeded in 39 out of 40 attempts, whilst Opus 5 achieved only minor improvements which also disrupted the application’s behaviour. In a separate test in which various Claude models were tasked with building a game from a single instruction, Opus 5.5 achieved the highest score for the quality of its graphics and finish.
It is worth noting the methodological caveat introduced by Anthropic itself: the results for OSWorld 2.0, Humanity’s Last Exam and Chartography correspond to the ‘with tools’ or partial mode, as the case may be, and not to unassisted execution. The figures from Terminal-Bench-Science 0.1 place GPT-6 Astra ahead of Opus 5.5 in that specific category, a fact that the official table clearly reflects.
3. Safety and Alignment Assessment
The section on safety forms a significant part of Anthropic’s message. Opus 5.5 has achieved the highest score to date in the company’s automated behavioural audit, a set of alignment tests that subjects the model to thousands of simulated scenarios. According to the published documentation, the model is less likely than recent versions to carry out actions that are difficult to reverse or to act outside the limits assigned to it, and is more resistant to instruction injection than Opus 5.
Anthropic has expanded the scope of its alignment tests to cover longer-duration tasks, impossible tasks and scenarios modelled on real-life incidents. The company explicitly acknowledges that the model ‘still has limitations’ and refers to the Opus 5.5 System Card for the full details of the evaluation, which have not been made public in the form of a summary of figures in the announcement.
Given its stated performance in biology and cybersecurity, which is comparable to that of Claude Mythos 5.1, Anthropic is rolling out Opus 5.5 with safeguards similar to those of Claude Fable 5.1. Verified organisations can apply for access from today via the Life Sciences Verification Programme for biological research. In the coming weeks, the company will expand the Cyber Verification Programme so that accredited cybersecurity professionals can use Opus 5.5 in their work.
4. Cost, Speed and Availability
The economic argument is clear: Opus 5.5 uses fewer computing resources than Opus 5, and its pricing reflects this. Anthropic estimates that typical workloads using the default configuration result in a 40 per cent saving. Input and output tokens are billed at $4 and $20 per million respectively, 20 per cent less than Opus 5. Cache reads, which the company says account for the bulk of the cost in agentic and coding work, fall to $0.20 per million tokens – a 60 per cent reduction.
| Price per million tokens | Claude Opus 5.5 | Claude Opus 5 |
|---|---|---|
| Cache read | $0.20 | $0.50 |
| Entry tokens | $4 | $5 |
| Output tokens | $20 | $25 |
| Cache writing | $5 | $6.25 |
In terms of speed, Anthropic claims that Opus 5.5 generates output more than 30 per cent faster than Opus 5. The company also announced two commercial changes to coincide with the launch: an increase in the five-hour usage limits for Pro, Max, Team and Enterprise plans with per-seat licensing, and the option for subscribers to save a speed limit reset to use at their convenience.
The model also incorporates a change in writing style. Early testers describe the writing as clearer and easier to follow than that of Opus 5, with the most relevant information placed at the beginning. Anthropic links this change to a security benefit as well as a practical one: a more organised text is easier to review and verify.
5. Impact on the Sector and Roadmap
This move reinforces Anthropic’s strategy of competing on efficiency rather than just raw capacity. The reduction in the cost of cache reads is significant for the segment that consumes the most resources: autonomous agents and code assistants, where the cumulative cost stems mainly from re-reading context. A 60 per cent reduction in this area has a direct impact on the cost-effectiveness of production deployments, rather than merely representing a one-off improvement in benchmark results.
The comparison published by Anthropic places Opus 5.5 ahead of GPT-6 Astra and GPT-5.6 Sol in most categories, although there are notable exceptions: GPT-6 Astra outperforms Opus 5.5 on AutomationBench (41.4% versus 40.0%) and on Terminal-Bench-Science 0.1 (64.6% versus 58.7%). The company does not hide these results in its own table, which reinforces the credibility of the rest of the figures.
The announced roadmap includes Claude Sonnet 5.5 and Claude Haiku 5.5 in the coming weeks, with corresponding improvements in performance, efficiency and security. In doing so, Anthropic is extending the capabilities of Opus 5.5 to the mid-range and low-end price points in its catalogue, which account for the bulk of usage in commercial applications.
Anthropic’s call to ‘set the pace at the frontier’, alongside the external assessment by METR and Frontier Design, sets out a distinctive position within the sector: the company presents a measured approach to roll-out as part of its value proposition. The announcement does not specify any additional transparency commitments regarding the System Card or any specific timelines beyond the release of the Sonnet and Haiku variants.
6. Conclusion
Claude Opus 5.5 is a release designed to reduce costs without compromising performance. Anthropic ranks it on a par with Fable 5.1 for most tasks, with a stated 40 per cent saving in execution costs and an improvement of over 30 per cent in generation speed. The published prices — $4 per million input tokens and $20 per million output tokens, with cache reads at $0.20 — are specific, verifiable figures set out in the official documentation.
For teams already running code agents or assistants in production, the most significant aspect of this announcement is not the improvement in benchmark results, but the reduction in cache read costs – precisely the cost that accounts for the lion’s share of the bill in such workloads. It is the combination of lower cost and faster output speed that is likely to drive adoption decisions in the short term.
There are two open questions. The first is to what extent Astra’s lead over GPT-6 holds up in real-world scenarios: Anthropic itself warns that, at this level of capability, benchmark margins are of little guidance. The second is to what extent the restricted-access safeguards will influence its adoption in cybersecurity and the life sciences, where the model requires prior verification by the company.
Español
English
Français
Português
Deutsch
Italiano