Skip to content
AI & Development AI Data

Claude Opus 5.5 Pricing and Performance Comparison

Listen or read this article

10:13

Listen, or read the text only.

Claude Opus 5.5 Pricing and Performance Comparison

Kokoro 82M AI-generated voice 11 min read

0:00 10:13

Advertisement

Download audio

File name
claude-opus-5-5-pricing-performance-comparison-en.mp3
Format
MP3 (audio/mpeg)
Duration
10:13
File size
7.02 MB
Engine
Kokoro 82M

This audio was generated by AI.

You may download and use it freely for personal use.

Claude Opus 5.5 Pricing and Performance Comparison

7 min read

Claude Opus 5.5 Pricing and Performance Comparison
Claude Opus 5.5 has cut its standard input and output token rates by 20% compared with Opus 5. The 40% reduction in task costs is based on Anthropic's evaluation; actual costs depend on token usage and cache use.
Claude Opus 5.5 standard rates are $4 per million input tokens and $20 per million output tokens.
The cache read rate is $0.20 per million tokens, 60% lower than Opus 5.
According to Anthropic, its Terminal-Bench 4.0 score rose from 52.3% to 66.4%.
The roughly 85% reduction in attempts to bypass boundaries comes from a specific safety evaluation.
To migrate an existing API, check the model name as well as inference and tool-calling settings.
Claude Opus 5.5 lowers standard token prices by 20%. Its coding and computer use evaluation scores have also improved. The 40% reduction in task costs is a result from Anthropic's testing.
Prices and evaluation figures are based on Anthropic's September 22, 2026 announcement.
Claude Opus 5.5 Price Comparison
Standard API prices are $4 per 1 million input tokens and $20 per 1 million output tokens. Tokens are the units a model uses to process text and other content. Input means content sent to the model, while output means content the model generates.
Billing item | Opus 5 | Opus 5.5 | Price reduction 1 million standard input tokens | $5 | $4 | 20% 1 million output tokens | $25 | $20 | 20% 1 million cache read tokens | $0.50 | $0.20 | 60%
Amounts are in US dollars. Caching stores inputs used repeatedly so they can be reused. The cost of creating a cache is calculated separately from the cost of reading it.
Conditions for 40% Savings on Task Costs
The 40% figure is not a discount that applies to every request. It is Anthropic's estimate of savings on typical tasks with default settings. It reflects both lower token prices and fewer tokens used per task.
Usage pattern | What affects cost | What to check Entering new content each time | Standard input and output usage | Actual token counts for both Reusing the same document or instructions | Cache creation and read usage | Whether the cache was actually reused A coding agent performing multiple steps | Repeated calls and retries | Total cost to complete the task Monthly subscription | Subscription price and usage limits | The distinction between token prices and subscription fees
Lower API prices alone do not mean lower monthly subscription fees. To determine the savings for your service, compare usage for the same tasks. Anthropic's Opus overview also describes the 40% figure as an estimate for token-billed tasks.
Calculation Example
Using 1 million input tokens and 1 million output tokens costs $24 at standard prices. For this comparison, assume usage is the same for both models. This calculation excludes caching and separate tool costs.
Calculation item | Opus 5 | Opus 5.5 1 million standard input tokens | $5 | $4 1 million output tokens | $25 | $20 Total | $30 | $24
The difference is $6, a 20% saving. This calculation does not include any reduction in tokens used per task. The result based on prices alone therefore differs from the announced 40%.
The price difference for reading 1 million cached tokens is $0.30. That amount compares only the read charge, however. The total bill must also account for cache creation and other token costs.
Agentic Coding and Computer Use Performance
In Anthropic's announcement, both coding and computer use scores improved. Agentic coding means using tools to carry out multistep development tasks. Computer use is the ability to view a screen and operate an interface.
Evaluation | Opus 5 | Opus 5.5 | Score difference Terminal-Bench 4.0 | 52.3% | 66.4% | 14.1 percentage points CursorBench 4.0 | 46.6% | 57.8% | 11.2 percentage points OSWorld 2.1 | 74.0% | 81.8% | 7.8 percentage points
The evaluation name in the official announcement is OSWorld 2.1. Its score is marked partial. Keep both the version and the scoring label to avoid confusing it with other results.
The Opus 5.5 result for Terminal-Bench 4.0 uses the xhigh setting. Its evaluation conditions are not the same as those used for the announcement's cost estimate with default settings. A higher score alone cannot tell you the success rate for a real project.
Changes in Speed and Writing
Anthropic announced that output generation is more than 30% faster than with Opus 5. This figure concerns the speed at which text is generated. Any reduction in total task time, including search and testing, must be measured separately.
Anthropic says writing has been improved to put the main point first and explain it clearly. It did not provide a separate improvement rate for the quality of Korean documents. Check the following with documents your team uses.
· Whether conclusions and supporting evidence are distinct · Whether the requested style and format are followed · Whether figures and quotations can be checked against the source · Whether the reasons for code changes are explained well enough to review
What the 85% Reduction in Safety Evaluations Means
The approximately 85% reduction refers to attempts to bypass isolation boundaries. The comparison is with Opus 5 or Claude Mythos 5.1. It does not mean that the rate of incidents in actual services fell by the same percentage.
A sandbox is an execution environment that limits a program's access. Model safety evaluations and that environment's access restrictions are separate matters. Better evaluation scores do not, by themselves, justify expanding operational permissions.
Anthropic explains that it is still difficult to detect every failure before deployment through evaluation. You can check the scope of the figure in the safety section of the Claude Opus 5.5 announcement. During operation, check both the scope of file access and execution logs.
Settings to Check When Switching APIs
Existing API integrations require checking settings beyond the model name. The Claude API model identifier is claude-opus-5-5. Other clouds use their respective platform model identifiers.
The official migration guide states the change to thinking settings as follows.
Thinking can't be disabled
The source is Anthropic's Migrating to Claude Opus 5.5. This means a setting to disable thinking is not supported.
· Check existing requests for settings that disable thinking. Opus 5.5 rejects those settings. · Specify effort, which controls thinking intensity. The default for Opus 5.5 is medium. · Check settings that force a specific tool call. Unsupported settings cause request errors. · Process response blocks according to their type. The first block is not always answer text. · Measure cost and completion time again using the same tasks. Thinking tokens are also included in output charges.
Support for AWS, Google Cloud, and Microsoft Azure was also announced at launch. Model identifiers and tool support conditions may differ by platform. Check the migration guide for the model used by your existing integration.
Common Mistakes
Read price and performance figures according to what they measure. Token prices, task costs, and evaluation scores are different metrics. Keeping the following distinctions in mind helps preserve the conditions for comparison.
Easy-to-confuse interpretation | Confirmed meaning Every bill falls by 40% | An estimate for typical tasks with default settings Caching cuts total costs by 60% | Only the cache read price falls by 60% Performance rises by 14.1% | The Terminal-Bench score rises by 14.1 percentage points Every task finishes 30% faster | Output generation speed improves by more than 30% Actual incidents fall by 85% | Boundary-bypass attempts in a specific evaluation fall by approximately 85%
Comparing Cost per Completed Task
When choosing a model, compare the cost per task that passes review. Token prices alone do not show the cost of retries and revisions. This is a way to apply the announced figures to actual work.
Calculate this by dividing the total execution cost by the number of tasks that pass review. Include the cost of failed attempts in the total execution cost. Recording human review time as a separate item also helps with comparison.
· Use the same task list and completion criteria · Separate standard input, cache, and output costs · Record total costs, including failures and retries · Record whether each task passes final review · Separate execution time from human revision time
0:00 0:00
1 / 49

Advertisement

Download text

File name
claude-opus-5-5-pricing-performance-comparison-en.txt
Format
TXT (text/plain)
Paragraphs
49

Downloads exactly what you see as a text file.

Please cite the source when quoting.

Large text

Makes the text larger and the colors clearer. Turn it on if the text feels too small.

Actual operating costs depend on token usage and cache use.

Key points

  • Claude Opus 5.5 standard rates are $4 per million input tokens and $20 per million output tokens.
  • The cache read rate is $0.20 per million tokens, 60% lower than Opus 5.
  • According to Anthropic, its Terminal-Bench 4.0 score rose from 52.3% to 66.4%.
  • The roughly 85% reduction in attempts to bypass boundaries comes from a specific safety evaluation.
  • To migrate an existing API, check the model name as well as inference and tool-calling settings.

Claude Opus 5.5 lowers standard token prices by 20%. Its coding and computer use evaluation scores have also improved. The 40% reduction in task costs is a result from Anthropic's testing.

Prices and evaluation figures are based on Anthropic's September 22, 2026 announcement.

Claude Opus 5.5 Price Comparison

Standard API prices are $4 per 1 million input tokens and $20 per 1 million output tokens. Tokens are the units a model uses to process text and other content. Input means content sent to the model, while output means content the model generates.

Billing item Opus 5 Opus 5.5 Price reduction
1 million standard input tokens $5 $4 20%
1 million output tokens $25 $20 20%
1 million cache read tokens $0.50 $0.20 60%

Amounts are in US dollars. Caching stores inputs used repeatedly so they can be reused. The cost of creating a cache is calculated separately from the cost of reading it.

When switching an existing API, check reasoning and tool-call settings as well as the model name.

Conditions for 40% Savings on Task Costs

The 40% figure is not a discount that applies to every request. It is Anthropic's estimate of savings on typical tasks with default settings. It reflects both lower token prices and fewer tokens used per task.

Usage pattern What affects cost What to check
Entering new content each time Standard input and output usage Actual token counts for both
Reusing the same document or instructions Cache creation and read usage Whether the cache was actually reused
A coding agent performing multiple steps Repeated calls and retries Total cost to complete the task
Monthly subscription Subscription price and usage limits The distinction between token prices and subscription fees

Lower API prices alone do not mean lower monthly subscription fees. To determine the savings for your service, compare usage for the same tasks. Anthropic's Opus overview also describes the 40% figure as an estimate for token-billed tasks.

Calculation Example

Using 1 million input tokens and 1 million output tokens costs $24 at standard prices. For this comparison, assume usage is the same for both models. This calculation excludes caching and separate tool costs.

Calculation item Opus 5 Opus 5.5
1 million standard input tokens $5 $4
1 million output tokens $25 $20
Total $30 $24

The difference is $6, a 20% saving. This calculation does not include any reduction in tokens used per task. The result based on prices alone therefore differs from the announced 40%.

The price difference for reading 1 million cached tokens is $0.30. That amount compares only the read charge, however. The total bill must also account for cache creation and other token costs.

Agentic Coding and Computer Use Performance

In Anthropic's announcement, both coding and computer use scores improved. Agentic coding means using tools to carry out multistep development tasks. Computer use is the ability to view a screen and operate an interface.

Evaluation Opus 5 Opus 5.5 Score difference
Terminal-Bench 4.0 52.3% 66.4% 14.1 percentage points
CursorBench 4.0 46.6% 57.8% 11.2 percentage points
OSWorld 2.1 74.0% 81.8% 7.8 percentage points

The evaluation name in the official announcement is OSWorld 2.1. Its score is marked partial. Keep both the version and the scoring label to avoid confusing it with other results.

The Opus 5.5 result for Terminal-Bench 4.0 uses the xhigh setting. Its evaluation conditions are not the same as those used for the announcement's cost estimate with default settings. A higher score alone cannot tell you the success rate for a real project.

Changes in Speed and Writing

Anthropic announced that output generation is more than 30% faster than with Opus 5. This figure concerns the speed at which text is generated. Any reduction in total task time, including search and testing, must be measured separately.

Anthropic says writing has been improved to put the main point first and explain it clearly. It did not provide a separate improvement rate for the quality of Korean documents. Check the following with documents your team uses.

  • Whether conclusions and supporting evidence are distinct
  • Whether the requested style and format are followed
  • Whether figures and quotations can be checked against the source
  • Whether the reasons for code changes are explained well enough to review

What the 85% Reduction in Safety Evaluations Means

The approximately 85% reduction refers to attempts to bypass isolation boundaries. The comparison is with Opus 5 or Claude Mythos 5.1. It does not mean that the rate of incidents in actual services fell by the same percentage.

A sandbox is an execution environment that limits a program's access. Model safety evaluations and that environment's access restrictions are separate matters. Better evaluation scores do not, by themselves, justify expanding operational permissions.

Anthropic explains that it is still difficult to detect every failure before deployment through evaluation. You can check the scope of the figure in the safety section of the Claude Opus 5.5 announcement. During operation, check both the scope of file access and execution logs.

Settings to Check When Switching APIs

Existing API integrations require checking settings beyond the model name. The Claude API model identifier is claude-opus-5-5. Other clouds use their respective platform model identifiers.

The official migration guide states the change to thinking settings as follows.

Thinking can't be disabled

The source is Anthropic's Migrating to Claude Opus 5.5. This means a setting to disable thinking is not supported.

  1. Check existing requests for settings that disable thinking. Opus 5.5 rejects those settings.
  2. Specify effort, which controls thinking intensity. The default for Opus 5.5 is medium.
  3. Check settings that force a specific tool call. Unsupported settings cause request errors.
  4. Process response blocks according to their type. The first block is not always answer text.
  5. Measure cost and completion time again using the same tasks. Thinking tokens are also included in output charges.

Support for AWS, Google Cloud, and Microsoft Azure was also announced at launch. Model identifiers and tool support conditions may differ by platform. Check the migration guide for the model used by your existing integration.

Common Mistakes

Read price and performance figures according to what they measure. Token prices, task costs, and evaluation scores are different metrics. Keeping the following distinctions in mind helps preserve the conditions for comparison.

Easy-to-confuse interpretation Confirmed meaning
Every bill falls by 40% An estimate for typical tasks with default settings
Caching cuts total costs by 60% Only the cache read price falls by 60%
Performance rises by 14.1% The Terminal-Bench score rises by 14.1 percentage points
Every task finishes 30% faster Output generation speed improves by more than 30%
Actual incidents fall by 85% Boundary-bypass attempts in a specific evaluation fall by approximately 85%

Comparing Cost per Completed Task

When choosing a model, compare the cost per task that passes review. Token prices alone do not show the cost of retries and revisions. This is a way to apply the announced figures to actual work.

Calculate this by dividing the total execution cost by the number of tasks that pass review. Include the cost of failed attempts in the total execution cost. Recording human review time as a separate item also helps with comparison.

  • Use the same task list and completion criteria
  • Separate standard input, cache, and output costs
  • Record total costs, including failures and retries
  • Record whether each task passes final review
  • Separate execution time from human revision time
Open the link

Opens in a new window.

Advertisement

Sign-in required

Sign in with your Google account to like, comment, and save highlights.

FAQ

When was Claude Opus 5.5 released?

Anthropic released Claude Opus 5.5 on September 22, 2026.

What is the standard API price for Claude Opus 5.5?

At release, the price was $4 per million input tokens and $20 per million output tokens. Each rate is 20% lower than for Opus 5.

Are costs always reduced by 40%?

The 40% figure is Anthropic's estimate for typical tasks. Actual savings vary depending on token usage and cache use.

How much did the cache read price drop?

It dropped from $0.50 to $0.20 per million tokens. That is a 60% reduction in the read rate. Cache creation costs need to be checked separately.

What is the exact name of the computer use evaluation?

The evaluation is called OSWorld 2.1 in the official announcement. Opus 5.5 scored 81.8%, with a partial label.

Does the 85% safety figure mean an 85% reduction in incidents?

It means that attempts to bypass isolation boundaries decreased by about 85% in a specific evaluation. It does not mean a reduction in incidents in actual services.

Can I just change the model name in my existing API requests?

Additional changes may be needed depending on your existing request settings. Check the migration documentation for reasoning settings, forced tool calls, and response block handling.

Is it available on AWS and Google Cloud too?

The launch announcement included support for AWS, Google Cloud, and Microsoft Azure. Check the model identifiers and conditions for feature support on each platform.

Do faster output generation and shorter task completion times mean the same thing?

They are different metrics. If a task includes tool execution and testing in addition to output generation, the total completion time may differ.

Sources

Also available as video and a short read

Videos and a short write-up made from this content. Watch instead of reading, or skim the gist first.

Short version

Loading…

Data formats

This content is available in several machine-friendly formats.

Data-only languages (machine translated, files only)

Indonesian JSON MD Portuguese JSON MD Chinese (Traditional) JSON MD Deutsch JSON MD

Verification

This article was drafted with AI and then reviewed and edited by a person.

Reviewed by 신익희 · 편집장 · 2026-10-02

Figures in this article were checked against the source material during generation. · 2026-10-02

This translation has been cross-checked by AI. · 2026-10-02

Reuse & AI usage

Search indexing and AI citation with attribution are welcome. See the license policy for details.

CC BY · License

Loading…

Loading…

From Injoys

Request the content you want and take 70% of what it earns

Just leave the subject. We handle production, review, translation and distribution.

See how revenue sharing works

Comments