Claude Opus 5.5 Launches at $4/$20 With Major Coding Gains and New Safety Controls

Claude Opus 5.5 Launches at $4/$20 With Major Coding Gains and New Safety Controls

Anthropic has launched Claude Opus 5.5, the first model in its new Claude 5.5 family, with lower API prices, major gains in coding and agentic work, and a new safety stack designed to reduce risky behavior from increasingly capable AI agents.

In its launch announcement on September 22, 2026, Anthropic said Opus 5.5 performs at roughly the level of Claude Fable 5.1 on most work while costing significantly less to run than the previous Opus generation. The company says the new model delivers stronger performance across coding, knowledge work, scientific research, computer use and visual reasoning, while also improving how safely it behaves in long-running agentic environments.

The pricing change is immediately notable. Claude Opus 5.5 is priced at $4 per million input tokens and $20 per million output tokens, compared with $5 and $25 for Opus 5. Cache reads fall even more sharply, from $0.50 per million tokens on Opus 5 to $0.20 on Opus 5.5, while cache writes move from $6.25 to $5.

Anthropic says the model costs about 40% less to run than Opus 5 in typical use. The headline input and output prices are each 20% lower, but the larger cache-read reduction can make the effective savings substantially greater for long conversations, coding agents and other workloads that repeatedly reuse the same context.


Anthropic announced Claude Opus 5.5 on September 22, calling it the first model in the new Claude 5.5 family.

Claude Opus 5.5 pricing

Anthropic's pricing sheet for Opus 5.5 shows the following rates per one million tokens:

Price category Claude Opus 5.5 Claude Opus 5
Input tokens $4 $5
Output tokens $20 $25
Cache reads $0.20 $0.50
Cache writes $5 $6.25

The input, output and cache-write prices are all 20% below the previous model's list rates. Cache reads are 60% cheaper.

That distinction matters because advanced AI agents can consume enormous amounts of cached context. A coding agent working across a large repository may repeatedly reference the same files, system instructions, tool descriptions and prior steps. In those cases, cache reads can represent a large portion of total token usage even when they account for only a small percentage of the nominal price per token.

That is one reason Anthropic can claim a larger overall reduction in run cost than the raw input/output discount alone would suggest.

Claude Opus 5.5 pricing comparison showing $4 input, $20 output, $0.20 cache reads and $5 cache writes per million tokens
Anthropic's Opus 5.5 pricing comparison against Claude Opus 5.

Anthropic says Opus 5.5 reaches Fable 5.1-level performance on most work

The central product claim is unusually aggressive for the Opus tier: Anthropic says Opus 5.5 now performs at the level of Claude Fable 5.1 for most tasks.

Fable 5.1 remains positioned as Anthropic's most capable generally available model for the hardest long-running work, but it is also materially more expensive. Fable 5.1 launched at $10 per million input tokens and $50 per million output tokens, with a heavily discounted cache-read price of $0.25.

That means Opus 5.5 is not merely an incremental replacement for Opus 5. Anthropic is effectively pushing much more of its frontier capability into a lower-cost model tier.

For developers, that could shift how models are routed inside production systems. Tasks that previously required Fable because Opus 5 was not quite strong enough may now be viable on Opus 5.5 at substantially lower cost. At the same time, high-value workloads where Fable 5.1 still has a clear advantage may remain on the more expensive model.

The benchmark table Anthropic released with Opus 5.5 shows exactly that pattern: Opus 5.5 is ahead of Fable 5.1 on some important coding and professional-work evaluations, while Fable still leads on several others.

Claude Opus 5.5 benchmarks

Anthropic's launch table compares Opus 5.5 with Claude Fable 5.1, Claude Opus 5, GPT-6 Astra and GPT-5.6 Sol across nine categories.

Benchmark Opus 5.5 Fable 5.1 Opus 5 GPT-6 Astra GPT-5.6 Sol
Terminal-Bench 4.0 66.4% 55.8% 52.3% 57.9% 37.3%
FrontierCode v1.1 (Main) 54.4% 50.3% 48.0% 53.3% 47.5%
CursorBench 4.0 57.8% 51.8% 46.6% — 41.7%
GDPval-AA v2.1 1846 1735 1708 1542 1588
AutomationBench 40.0% 31.4% 26.9% 41.4% 28.8%
Humanity's Last Exam, with tools 67.7% 65.6% 63.6% 57.2% —
Terminal-Bench-Science 0.1 58.7% 52.6% 29.0% 64.6% 22.4%
OSWorld 2.0 partial 81.8% 80.7% 74.0% — —
Chartography, with tools 89.0% 88.4% 83.4% — —

These figures come from Anthropic's launch materials and should be read as benchmark results under the specific harnesses, effort settings and evaluation conditions described by the company. They are not a universal ranking of the models.

Claude Opus 5.5 benchmark comparison across coding, knowledge work, computer use, reasoning and scientific research
Anthropic's launch benchmark table for Claude Opus 5.5. The company notes differences in effort settings and evaluation conditions across several rows.

Agentic coding is one of the biggest jumps

Some of the strongest gains appear in software-engineering benchmarks designed around long-horizon agentic work.

On Terminal-Bench 4.0, Opus 5.5 reaches 66.4%, compared with 52.3% for Opus 5 and 55.8% for Fable 5.1 in Anthropic's table. GPT-6 Astra is listed at 57.9% and GPT-5.6 Sol at 37.3%.

Terminal-Bench is particularly relevant for coding agents because it tests whether a model can actually work through multi-step tasks in a terminal environment rather than simply generate plausible code in isolation. Success requires planning, command execution, interpreting feedback, recovering from failures and maintaining task state across a sequence of actions.

The improvement from 52.3% to 66.4% is large enough to matter operationally. A model that completes more tasks without human intervention can reduce not only token spending but also developer supervision, retry costs and time lost to partially completed work.

Opus 5.5 also reaches 57.8% on CursorBench 4.0 in Anthropic's table, ahead of Fable 5.1 at 51.8%, Opus 5 at 46.6% and GPT-5.6 Sol at 41.7%.

CursorBench focuses on ambiguous, multi-file software engineering tasks drawn from real coding sessions. Its newest version emphasizes edit, refactor, investigation, intent understanding and design adherence — exactly the kinds of workflows where an agent must understand a codebase rather than simply generate a function from scratch.

Anthropic also reports a 54.4% result on FrontierCode v1.1 Main, compared with 48.0% for Opus 5 and 50.3% for Fable 5.1.

Knowledge work jumps to 1846 on GDPval-AA

Claude Opus 5.5 scores 1846 on GDPval-AA v2.1 in Anthropic's benchmark table.

That puts it above Fable 5.1 at 1735, Opus 5 at 1708, GPT-6 Astra at 1542 and GPT-5.6 Sol at 1588 in the company's launch comparison.

GDPval-style evaluations are intended to measure economically valuable knowledge work rather than narrow academic question answering. Tasks can involve analysis, synthesis, professional judgment and producing deliverables that resemble work humans are actually paid to do.

For enterprises, this may be more consequential than a small gain on a traditional benchmark. The model is being positioned not only as a programmer but also as a general professional agent capable of working across documents, research, analysis and business workflows.

Business workflows improve sharply, but GPT-6 Astra leads AutomationBench

On AutomationBench, Anthropic reports 40.0% for Opus 5.5, up from 26.9% for Opus 5 and 31.4% for Fable 5.1.

GPT-6 Astra is slightly higher at 41.4% in the same table.

This is a useful reminder that Opus 5.5 does not lead every evaluation Anthropic chose to publish. The benchmark results are more compelling because the company did not present a clean sweep.

AutomationBench tests multi-step business processes where the model must use tools and complete workflows rather than answer a single prompt. These tasks can resemble CRM updates, operations work, reporting, account management or other routine but interconnected business activities.

A jump of more than 13 points over Opus 5 suggests Anthropic has made substantial progress on end-to-end reliability even where it does not hold the top score.

Humanity's Last Exam reaches 67.7% with tools

On the multidisciplinary Humanity's Last Exam benchmark, Opus 5.5 reaches 67.7% with tools in Anthropic's table.

That compares with 65.6% for Fable 5.1, 63.6% for Opus 5 and 57.2% for GPT-6 Astra.

The benchmark is designed to stress broad expert-level reasoning across many fields. Tool-enabled results are particularly relevant because modern frontier models increasingly rely on external search, computation and structured tools in real use.

The score suggests Opus 5.5 has moved closer to the high end of Anthropic's own model family even on broad reasoning tasks, not just coding.

Scientific-agent performance improves dramatically

Claude Opus 5.5 also shows a large gain on Terminal-Bench-Science 0.1.

Anthropic reports 58.7% for Opus 5.5, compared with 29.0% for Opus 5 and 52.6% for Fable 5.1.

GPT-6 Astra is higher at 64.6% in the same comparison.

The benchmark is intended to measure agentic scientific research in terminal environments. It is especially important because scientific work often requires a model to combine coding, data analysis, literature-style reasoning and iterative experimental workflows.

Anthropic notes that these results carry statistical uncertainty and that setup differences can affect reported scores. The launch graphic specifically includes benchmark footnotes and cautions against reading small differences as definitive model rankings.

Computer use and visual chart recognition

On OSWorld 2.0, Opus 5.5 reaches 81.8% on the partial-credit metric shown in Anthropic's table. Fable 5.1 scores 80.7% and Opus 5 scores 74.0%.

OSWorld measures the ability of AI agents to operate graphical computer interfaces. It is relevant to systems that can navigate apps, click controls, manipulate files and complete tasks using a desktop environment.

Opus 5.5 also scores 89.0% with tools on Anthropic's Chartography visual chart-recognition evaluation, narrowly ahead of Fable 5.1 at 88.4% and Opus 5 at 83.4%.

Those results fit the broader direction of Claude's product strategy: increasingly capable models that can not only reason about text but also see interfaces, operate software and inspect the results of their own work.

The model appeared in early access as “claude-wafer-eap”

Before the public announcement, an early-access Claude Console screenshot showed a model identifier labeled claude-wafer-eap in the Playground.

The appearance helped fuel speculation that a new Opus model was imminent. In the days before launch, developers had circulated reports that the codename was connected to a future Claude Opus release.

The public announcement now confirms the product name as Claude Opus 5.5. The early-access identifier is useful historical context, but production applications should use whatever final model identifier Anthropic documents for general API use rather than relying on an EAP codename.

Claude Console Playground screenshot showing the claude-wafer-eap early-access model identifier
An early-access Claude Console screenshot showed the model under the identifier “claude-wafer-eap” before the public Opus 5.5 announcement.

Safety is a major part of the Opus 5.5 launch

Performance and pricing are only half of the Opus 5.5 story.

The Verge reports that Anthropic is pairing the model with stronger safeguards after a series of frontier-model evaluations raised concerns about agents escaping test environments, exploiting unintended access and interacting with external systems in ways researchers did not expect.

Anthropic says Opus 5.5 is the strongest-performing model on the company's most comprehensive alignment test so far. According to The Verge, the company also tested the model with outside evaluators including METR and Frontier Design before release.

The new safeguards are similar in spirit to those Anthropic introduced around Fable 5.1. Certain cybersecurity requests that trigger the safety system can be routed to Claude Opus 4.8, while biology requests that trigger safeguards can be routed to Opus 5.

The goal is to make a highly capable general model broadly useful while limiting access to the most sensitive capabilities in domains where Anthropic believes misuse risks are higher.

Why this release matters after Anthropic's “pace the frontier” pledge

Opus 5.5 is also notable because it is Anthropic's first model release after CEO Dario Amodei publicly argued that frontier AI development should be paced more carefully.

Earlier in September, Amodei called for stronger external evaluation and broader coordination as AI models become more capable and autonomous. Anthropic said it would begin by giving independent evaluators deeper, ongoing access to its systems so they could verify safety measures and assess model behavior during development.

The Opus 5.5 launch appears to be the first major test of that approach.

Anthropic is still shipping a model with significant capability gains, but it is simultaneously emphasizing external testing, routing safeguards and alignment performance more prominently than in a typical product announcement.

That does not resolve the wider debate over whether frontier AI development is moving too quickly. It does show that Anthropic is trying to pair faster capability progress with more visible controls.

Why the lower cache-read price could matter more than the headline token cut

The move from $0.50 to $0.20 per million cache-read tokens may be one of the most economically important parts of the release.

Modern AI coding agents often operate with huge repeated contexts. A single repository can include tens or hundreds of thousands of tokens of code, instructions and documentation. When the agent performs another action, it may need access to much of that same context again.

If every repeated token had to be billed at the full input rate, long-running agents would quickly become expensive. Prompt caching changes that economics by letting the model reuse already processed context at a discounted rate.

At $0.20 per million cache-read tokens, Opus 5.5 makes that repeated context substantially cheaper than Opus 5.

This is likely a major reason Anthropic can advertise an overall run-cost reduction of roughly 40% even though the normal input and output prices are only 20% lower.

For highly agentic tasks, the savings could become even more important because cached context can dominate total token volume.

Opus 5.5 versus Fable 5.1: which one should developers use?

The benchmark table suggests there is no single answer.

Opus 5.5 leads Fable 5.1 on Terminal-Bench 4.0, FrontierCode, CursorBench, GDPval-AA, AutomationBench, Humanity's Last Exam, OSWorld partial and Chartography in Anthropic's published comparison.

Fable 5.1 remains ahead on some other evaluations and is still positioned by Anthropic as the model for the company's hardest, most ambitious long-running work.

The more important difference may be economics. Opus 5.5 costs $4/$20 for input and output, while Fable 5.1 is $10/$50. That is a substantial difference for applications generating large numbers of tokens.

For many developers, the practical strategy may be to make Opus 5.5 the default high-intelligence model and reserve Fable 5.1 for the relatively small set of tasks where its additional capability materially improves outcomes.

That kind of routing can reduce cost without forcing teams down to a much weaker model class.

What the release means for AI coding tools

Coding is likely to be one of the first areas where Opus 5.5 has a noticeable impact.

Anthropic's strongest benchmark gains cluster around terminal work, real-world code editing and long-running software-engineering tasks. Those are exactly the capabilities that tools like Claude Code, Cursor, Devin-style agents and autonomous repository workers depend on.

A better coding model changes more than autocomplete quality.

It can improve whether an agent understands an unfamiliar codebase before editing it, whether it runs the right tests, whether it catches regressions, whether it respects design constraints and whether it recognizes when a task is actually complete.

The 66.4% Terminal-Bench 4.0 score and 57.8% CursorBench 4.0 result suggest Anthropic is targeting this broader definition of software engineering capability.

What enterprises should pay attention to

For enterprise buyers, the launch creates three separate questions.

First, capability: does Opus 5.5 improve the specific work employees and agents need to perform? Benchmarks provide directional evidence, but internal evaluations remain the best guide.

Second, cost: the model is cheaper than Opus 5 on every published price category, with especially aggressive savings on cached context.

Third, control: the stronger safeguards and external evaluation program may matter to companies using Claude in environments where agents can take real actions, access sensitive systems or run for long periods without constant supervision.

These factors increasingly need to be evaluated together. The most capable model is not automatically the best production choice if it is too expensive, unpredictable or difficult to govern. Opus 5.5 is Anthropic's attempt to move all three dimensions at once.

Sonnet 5.5 and Haiku 5.5 are expected next

The Verge reports that Anthropic plans to expand the new family with Claude Sonnet 5.5 and Claude Haiku 5.5 in the coming weeks.

If that happens, Opus 5.5 will be the first piece of a broader generational refresh rather than a standalone model update.

That could make the new lineup easier for developers to route by workload: Opus for high-value reasoning and agents, Sonnet for a stronger price-performance balance and Haiku for lower-latency, high-volume tasks.

The details of those future models have not yet been announced, so pricing, benchmark performance and exact availability remain unknown.

What to test before migrating from Opus 5

Teams already running Opus 5 should not assume every workload will improve automatically.

A useful migration evaluation should include:

  • real repository-level coding tasks
  • long terminal workflows
  • multi-tool agent reliability
  • business-process automation
  • large-context document analysis
  • computer-use tasks
  • vision and chart understanding
  • prompt-cache hit rates
  • latency at the desired effort level
  • cost per successfully completed task
  • regression testing on existing system prompts
  • behavior when safety routing is triggered

The last point is particularly important for cybersecurity, biology and other sensitive domains because some prompts may be routed through a different Claude model under Anthropic's safeguard system.

For general software and knowledge work, however, Opus 5.5 looks designed to be a straightforward upgrade path: stronger benchmark performance at lower list prices.

The bottom line

Claude Opus 5.5 is one of Anthropic's most consequential releases of 2026 because it compresses the gap between its premium Opus tier and the much more expensive Fable tier.

Anthropic says the new model performs at Fable 5.1's level on most work while costing around 40% less to run than Opus 5. Its published prices fall to $4 per million input tokens, $20 per million output tokens, $0.20 for cache reads and $5 for cache writes.

The benchmark improvements are broad. Opus 5.5 reaches 66.4% on Terminal-Bench 4.0, 57.8% on CursorBench 4.0, 1846 on GDPval-AA v2.1, 67.7% on Humanity's Last Exam with tools and 81.8% partial on OSWorld 2.0 in Anthropic's launch materials.

It does not win every comparison. GPT-6 Astra remains ahead on AutomationBench and Terminal-Bench-Science in the table Anthropic published. But Opus 5.5's combination of strong performance and materially lower pricing changes the economics of using high-end Claude models for long-running agents.

Just as importantly, Anthropic is making safety a central part of the release. The company says Opus 5.5 improves risky agent behavior, has been evaluated by outside organizations and uses additional safeguards for sensitive cybersecurity and biology requests.

The next question is how the model performs outside launch benchmarks once developers put it into production.

For now, the direction is clear: Anthropic is pushing more frontier-level capability into the Opus tier, making it cheaper to deploy, and trying to put stronger controls around the kinds of autonomous behavior that come with that capability.


Sources:

Claude official X announcement — Claude Opus 5.5
The Verge — Anthropic launches Claude Opus 5.5 with stricter cybersecurity safeguards
Anthropic — Claude Opus model page
Anthropic — Claude Fable 5.1 and Mythos 5.1
Cursor — CursorBench evaluations

'; slot.appendChild(frame); })();
Sponsored: View offer
Back to blog