Claude Opus 5.5 is Anthropic's new high-end model for agentic coding and professional work, launched with a claim that typical workloads cost 40% less than on Opus 5. The September 22 release also raises the company's benchmark scores while extending safeguards for sensitive cyber and biology tasks.

 

The release changes four practical areas for developers and enterprises:

  • Input and output prices fall to $4 and $20 per million tokens.
  • Default workloads use fewer tokens and run more than 30% faster.
  • Agentic coding and knowledge-work scores improve across several tests.
  • Advanced cyber and biology access remains restricted to vetted users.

 

Related Coverage

 

Claude Opus 5.5 Pricing Cuts

Anthropic lists standard API pricing at $4 per million input tokens and $20 per million output tokens, down 20% from Opus 5. Cache reads fall from $0.50 to $0.20 per million tokens, a 60% reduction aimed squarely at coding agents that repeatedly consult stored context.

 

The company's 40% savings figure is broader than the per-token discount. Anthropic says Opus 5.5 also needs fewer tokens to finish typical tasks, so the total bill can fall faster than the list price alone suggests. Output generation is more than 30% faster, according to the launch materials.

 

A faster mode is available through Claude Code and the Claude Platform at up to 2.5 times standard speed. It costs $8 per million input tokens and $40 per million output tokens, preserving the familiar trade-off between latency and price.

 

Benchmark Gains and Their Limits

On Anthropic's reported results, Opus 5.5 scored 66.4% on Terminal-Bench 4.0, compared with 52.3% for Opus 5. It reached 54.4% on the main FrontierCode v1.1 evaluation and 57.8% on CursorBench 4.0, tests focused on multi-step software engineering.

 

The model also posted an 1846 Elo score on GDPval-AA v2.1, which evaluates professional work across 44 occupations. On Humanity's Last Exam with tools, Anthropic reported 67.7%, ahead of the company's Fable 5.1 and Opus 5 results.

 

Those comparisons require care. Anthropic notes that models were sometimes run at different reasoning-effort settings and that several competitor scores were reported by their makers. It also warns that narrow benchmark gaps are becoming less reliable guides to real-world differences among frontier systems.

 

Long-Running Coding and Knowledge Work

Opus 5.5 is designed for jobs that keep an agent working across many files, tools, and decisions. Anthropic says an early tester used it to audit and repair a 200,000-line codebase in under three hours, versus more than 20 hours for Opus 5 with 2.5 times as many tokens.

 

In an internal translation experiment, Opus 5.5 and Fable 5.1 rewrote the HAProxy load balancer from C to Rust. Both versions passed nearly all of HAProxy's regression tests, but Opus 5.5 finished in 9.5 hours, three hours sooner, at 51% lower cost.

 

The model's efficiency could matter as software teams shift from short chat exchanges to agents that work unattended. Token price remains important, but tool calls, retries, context reads, and the number of steps needed to finish a task increasingly determine the real operating cost.

 

Knowledge-work tests point in the same direction. Anthropic reports that 16 of 18 Opus 5.5 research reports passed an automated factuality threshold in an internal evaluation where a single invented figure or quotation caused failure. Neither Fable 5.1 nor Opus 5 passed under the same rule.

 

Safeguards, Access and Migration Changes

Anthropic describes Opus 5.5 as its strongest model yet on an automated behavioral audit covering thousands of simulated scenarios. The company says it is less likely than recent models to take irreversible actions outside instructions and is more resistant than Opus 5 to prompt-injection attacks.

 

Because the model approaches Anthropic's more restricted systems in cybersecurity and biology, general users do not receive unrestricted access to every capability. Most cyber tasks are rerouted to Opus 4.8, while vetted researchers can apply to verification programs for broader cyber and life-sciences use.

 

Developers upgrading from Opus 5 must also account for compatibility changes. Thinking cannot be disabled, forced tool use can return an error, thinking blocks are tied to the model and conversation, and an older computer-use tool version is no longer accepted on the Claude API or Google Cloud.

 

Opus 5.5 is available through Claude for Pro, Max, Team, and Enterprise customers, as well as the Claude API and major cloud platforms. Anthropic says Sonnet 5.5 and Haiku 5.5 will follow, making the launch the first step in a wider family refresh rather than a one-off model update.