Claude Opus 5.5 Pricing and Performance Comparison
Claude Opus 5.5 has cut its standard input and output token rates by 20% compared with Opus 5. The 40% reduction in task costs is based on Anthropic's evaluation; actual costs depend on token usage and cache use.
Claude Opus 5.5 standard rates are $4 per million input tokens and $20 per million output tokens.
The cache read rate is $0.20 per million tokens, 60% lower than Opus 5.
According to Anthropic, its Terminal-Bench 4.0 score rose from 52.3% to 66.4%.
The roughly 85% reduction in attempts to bypass boundaries comes from a specific safety evaluation.
To migrate an existing API, check the model name as well as inference and tool-calling settings.
Claude Opus 5.5 lowers standard token prices by 20%. Its coding and computer use evaluation scores have also improved. The 40% reduction in task costs is a result from Anthropic's testing.
Prices and evaluation figures are based on Anthropic's September 22, 2026 announcement.
Claude Opus 5.5 Price Comparison
Standard API prices are $4 per 1 million input tokens and $20 per 1 million output tokens. Tokens are the units a model uses to process text and other content. Input means content sent to the model, while output means content the model generates.
Billing item | Opus 5 | Opus 5.5 | Price reduction
1 million standard input tokens | $5 | $4 | 20%
1 million output tokens | $25 | $20 | 20%
1 million cache read tokens | $0.50 | $0.20 | 60%
Amounts are in US dollars. Caching stores inputs used repeatedly so they can be reused. The cost of creating a cache is calculated separately from the cost of reading it.
Conditions for 40% Savings on Task Costs
The 40% figure is not a discount that applies to every request. It is Anthropic's estimate of savings on typical tasks with default settings. It reflects both lower token prices and fewer tokens used per task.
Usage pattern | What affects cost | What to check
Entering new content each time | Standard input and output usage | Actual token counts for both
Reusing the same document or instructions | Cache creation and read usage | Whether the cache was actually reused
A coding agent performing multiple steps | Repeated calls and retries | Total cost to complete the task
Monthly subscription | Subscription price and usage limits | The distinction between token prices and subscription fees
Lower API prices alone do not mean lower monthly subscription fees. To determine the savings for your service, compare usage for the same tasks. Anthropic's Opus overview also describes the 40% figure as an estimate for token-billed tasks.
Calculation Example
Using 1 million input tokens and 1 million output tokens costs $24 at standard prices. For this comparison, assume usage is the same for both models. This calculation excludes caching and separate tool costs.
Calculation item | Opus 5 | Opus 5.5
1 million standard input tokens | $5 | $4
1 million output tokens | $25 | $20
Total | $30 | $24
The difference is $6, a 20% saving. This calculation does not include any reduction in tokens used per task. The result based on prices alone therefore differs from the announced 40%.
The price difference for reading 1 million cached tokens is $0.30. That amount compares only the read charge, however. The total bill must also account for cache creation and other token costs.
Agentic Coding and Computer Use Performance
In Anthropic's announcement, both coding and computer use scores improved. Agentic coding means using tools to carry out multistep development tasks. Computer use is the ability to view a screen and operate an interface.
Evaluation | Opus 5 | Opus 5.5 | Score difference
Terminal-Bench 4.0 | 52.3% | 66.4% | 14.1 percentage points
CursorBench 4.0 | 46.6% | 57.8% | 11.2 percentage points
OSWorld 2.1 | 74.0% | 81.8% | 7.8 percentage points
The evaluation name in the official announcement is OSWorld 2.1. Its score is marked partial. Keep both the version and the scoring label to avoid confusing it with other results.
The Opus 5.5 result for Terminal-Bench 4.0 uses the xhigh setting. Its evaluation conditions are not the same as those used for the announcement's cost estimate with default settings. A higher score alone cannot tell you the success rate for a real project.
Changes in Speed and Writing
Anthropic announced that output generation is more than 30% faster than with Opus 5. This figure concerns the speed at which text is generated. Any reduction in total task time, including search and testing, must be measured separately.
Anthropic says writing has been improved to put the main point first and explain it clearly. It did not provide a separate improvement rate for the quality of Korean documents. Check the following with documents your team uses.
· Whether conclusions and supporting evidence are distinct
· Whether the requested style and format are followed
· Whether figures and quotations can be checked against the source
· Whether the reasons for code changes are explained well enough to review
What the 85% Reduction in Safety Evaluations Means
The approximately 85% reduction refers to attempts to bypass isolation boundaries. The comparison is with Opus 5 or Claude Mythos 5.1. It does not mean that the rate of incidents in actual services fell by the same percentage.
A sandbox is an execution environment that limits a program's access. Model safety evaluations and that environment's access restrictions are separate matters. Better evaluation scores do not, by themselves, justify expanding operational permissions.
Anthropic explains that it is still difficult to detect every failure before deployment through evaluation. You can check the scope of the figure in the safety section of the Claude Opus 5.5 announcement. During operation, check both the scope of file access and execution logs.
Settings to Check When Switching APIs
Existing API integrations require checking settings beyond the model name. The Claude API model identifier is claude-opus-5-5. Other clouds use their respective platform model identifiers.
The official migration guide states the change to thinking settings as follows.
Thinking can't be disabled
The source is Anthropic's Migrating to Claude Opus 5.5. This means a setting to disable thinking is not supported.
· Check existing requests for settings that disable thinking. Opus 5.5 rejects those settings.
· Specify effort, which controls thinking intensity. The default for Opus 5.5 is medium.
· Check settings that force a specific tool call. Unsupported settings cause request errors.
· Process response blocks according to their type. The first block is not always answer text.
· Measure cost and completion time again using the same tasks. Thinking tokens are also included in output charges.
Support for AWS, Google Cloud, and Microsoft Azure was also announced at launch. Model identifiers and tool support conditions may differ by platform. Check the migration guide for the model used by your existing integration.
Common Mistakes
Read price and performance figures according to what they measure. Token prices, task costs, and evaluation scores are different metrics. Keeping the following distinctions in mind helps preserve the conditions for comparison.
Easy-to-confuse interpretation | Confirmed meaning
Every bill falls by 40% | An estimate for typical tasks with default settings
Caching cuts total costs by 60% | Only the cache read price falls by 60%
Performance rises by 14.1% | The Terminal-Bench score rises by 14.1 percentage points
Every task finishes 30% faster | Output generation speed improves by more than 30%
Actual incidents fall by 85% | Boundary-bypass attempts in a specific evaluation fall by approximately 85%
Comparing Cost per Completed Task
When choosing a model, compare the cost per task that passes review. Token prices alone do not show the cost of retries and revisions. This is a way to apply the announced figures to actual work.
Calculate this by dividing the total execution cost by the number of tasks that pass review. Include the cost of failed attempts in the total execution cost. Recording human review time as a separate item also helps with comparison.
· Use the same task list and completion criteria
· Separate standard input, cache, and output costs
· Record total costs, including failures and retries
· Record whether each task passes final review
· Separate execution time from human revision time