This is the English edition. 한국어판 and 日本語版 are also available.

Claude Opus 5.5 Performance Review: Benchmarks, Price and Who Should Use It

2026-09-23 · Updated 2026-09-25 · AI · United States · Zoogom Editorial

#Claude Opus 5.5#Anthropic#AI benchmarks#Claude Code#GPT-6 Astra#generative AI

A U.S. engineering team evaluating a frontier AI system across complex workstreams

Claude Opus 5.5 officially launched on September 22, 2026. What had been an unverified product name and price claim is now documented in Anthropic’s launch announcement, its API catalog and independent benchmark results.

The short verdict is more nuanced than “new benchmark champion.” Opus 5.5 is a top-tier model for long-running coding agents and professional knowledge work, while its lower token prices and cheaper prompt-cache reads make those sessions less expensive than Opus 5. Maximum effort delivers the highest scores, but it can also generate far more reasoning and output tokens. For many production workloads, the default medium setting is the more important result.

Our earlier rumor audit separated confirmed information from claims circulating before launch. This article focuses only on the specifications and measurements now available.

Core specifications and price

Core specifications and price: Item, Official Opus 5.5 figure, What it means in practice

The Claude Platform model page lists a 1 million-token context window, 128,000-token standard maximum output and a 300,000-token beta maximum for the Batch API. Both the reliable knowledge cutoff and training-data cutoff are June 2026.

First-party API pricing is $4 per million input tokens and $20 per million output tokens, down 20% from Opus 5 at $5 and $25. A five-minute cache write costs $5, a one-hour cache write costs $8 and a cache read costs $0.20 per million tokens. Batch processing halves base input and output prices to $2 and $10.

Anthropic says the model also uses fewer tokens and less serving compute on typical workloads, producing an estimated 40% reduction in task cost at default settings. Cache reads fell 60% from Opus 5’s $0.50 rate. That cache change can matter more than the headline input price for coding agents that repeatedly reuse the same repository and conversation state.

Anthropic benchmarks: strengths and limits

Anthropic benchmark strengths and limits: Evaluation, Opus 5.5, Comparison shown by Anthropic

Anthropic’s reported results favor Opus 5.5 on terminal-based agentic coding, merge-oriented code changes and professional knowledge work across 44 occupations. Against Opus 5, Terminal-Bench 4.0 rises from 52.3% to 66.4%, while GDPval-AA moves from 1708 to 1846 Elo.

The same table prevents a universal winner claim. GPT-6 Astra reaches 41.4% on AutomationBench against Opus 5.5’s 40.0%. Astra also leads Terminal-Bench-Science at 64.6% versus 58.7%. The better model depends on whether the workload resembles repository work, connected-app automation, scientific terminal tasks or polished business deliverables.

Effort settings also differ. Most reported Opus 5.5 scores use max, while Terminal-Bench uses xhigh; Astra figures use the strongest settings reported by OpenAI for each test. A one-point difference is not a controlled everyday comparison. Anthropic itself cautions that real-world differences between frontier models can be narrower than benchmark margins suggest.

Five increasing reasoning paths visualizing the tradeoff between effort, time and finished output

Independent results: the top score comes with high token use

Artificial Analysis independently measured Opus 5.5 at 58 on its Intelligence Index at maximum effort, the highest launch-day score in its dataset. The model led six of the ten component evaluations, including Humanity’s Last Exam, SciCode, GDPval-AA, AA-Briefcase, AA-Omniscience and AutomationBench-AA.

Its AA-Briefcase score reached 1822 Elo, 143 points above Claude Fable 5.1, with gains in both analysis and presentation. Terminal-Bench 4.0 reached 59.6%, effectively level with GPT-6 Astra in that independent setup. The external results therefore support the general direction of Anthropic’s claim: agentic coding and professional work are genuine strengths.

Maximum effort is not a free upgrade, however. Artificial Analysis measured roughly 119,000 output tokens per Intelligence Index task for Opus 5.5 max, compared with about 73,000 for Opus 5, 78,000 for Fable 5.1 and 27,000 for GPT-6 Astra max. Lower per-token pricing kept task cost near Opus 5 despite roughly 1.6 times the output, but long responses and waiting time remain operational costs.

The medium setting scored 51 in one Artificial Analysis run at roughly $1.34 per task. xhigh reached 56 at about $3.46 per task. Those numbers can change as providers and benchmarks update, but the practical message is stable: start with medium and raise effort only when higher accuracy reduces enough failures to pay for itself.

How it compares with GPT-6 Astra and Sol

Cross-company claims are most useful when the same evaluator runs both models. In an Artificial Analysis comparison, Opus 5.5 high scored 54 on the Intelligence Index while GPT-6 Astra max scored 53. Opus 5.5 had an advantage in output speed and its blended token price, while Astra retained strengths in some automation and scientific tasks and could have lower initial latency depending on effort.

Against GPT-6 Sol max, Opus 5.5 high scored 54 to 48 and reached 57% to Sol’s 44% on Terminal-Bench 4.0. Sol remained less expensive per token and produced output faster. A sensible U.S. deployment may route complex repository changes and high-stakes analysis to Opus, then use Sol or another fast model for high-volume classification, extraction and routine transformations.

Opus 5.5 and Astra are listed with 1 million-token context windows, while Sol is listed at 870,000 tokens. Context capacity does not guarantee reliable recall across the entire window. Teams should test middle-of-document retrieval, cross-file conflicts, citation accuracy and omission rates on their own material before treating the headline limit as usable memory.

Which effort level should you use?

Adaptive thinking is always enabled on Opus 5.5; applications cannot turn it off. Running every request at max can turn a modest quality gain into much higher token use and slower completion. A production router should begin at medium, escalate known hard tasks and keep simpler workloads on cheaper models.

What changed for long-running coding agents

Anthropic describes an early test in which Opus 5.5 audited and repaired a 200,000-line codebase in under three hours. Opus 5 reportedly took more than 20 hours and used 2.5 times as many tokens. In an internal C-to-Rust HAProxy migration, Opus 5.5 finished in 9.5 hours compared with 12 hours for Fable 5.1 and cost 51% less.

These are vendor and early-tester cases, not guarantees for every repository. They still reveal the design target: preserving a plan across many steps, tracking dependencies between repositories, repeatedly testing changes and checking work without constant human prompting. A serious pilot should measure regression pass rate, files changed, reverted commits, human correction time and total completion cost—not only whether an agent eventually produced code.

Migration issues developers should not miss

The Opus 5.5 migration notes list several breaking or visible behavior changes.

Changing the model ID to claude-opus-5-5 is only the first step. Production systems should test refusals, fallback routing, tool schemas, prompt-cache behavior and streaming presentation before switching live traffic.

Post-launch safety review and the Pace the Frontier context

Opus 5.5 is Anthropic’s first major model after Dario Amodei published We Must Pace the Frontier. That sequence invites a fair question: how can a company call for slower frontier progress and release a more capable system ten days later? In Amodei’s framework, pacing is not a total development stop. It means capability growth should be conditioned on safety, alignment, interpretability and third-party verification keeping up.

Anthropic says Frontier Design and METR participated in pre-release external evaluation. Product safeguards include classifiers before consequential actions, auditable sandboxes, a Life Sciences Verification Program and a Cyber Verification Program. These measures are evidence supporting Anthropic’s safety case, but they are not the same as a fully public independent audit that discloses scope, duration, unresolved findings and an evaluator’s authority to delay release.

Preserved thinking can improve continuity in long agent runs, but it also makes integration discipline important. Thinking blocks are tied to the model and conversation; rewriting or reusing them incorrectly can break compatibility. Anthropic also says it added defenses against distillation attacks that try to copy model capabilities from outputs. Developers should follow thinking-block storage rules and set their own retention policy for sensitive traces.

Anthropic says Sonnet 5.5 and Haiku 5.5 will follow within weeks. The current routing and price conclusions are therefore not permanent. Opus may remain the choice for high-failure-cost complex work, while the next Sonnet and Haiku releases could take more high-volume and latency-sensitive traffic. Teams should rerun the same completed-task-cost evaluation when those models arrive.

What U.S. customers should know

Opus 5.5 is available through the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Anthropic’s AWS platform. Anthropic also says it increased five-hour usage limits for Pro, Max, Team and seat-based Enterprise plans and provided subscription users with a reset they can save and use later. Account-level availability and practical limits can still vary by plan.

API billing and Claude subscriptions remain different cost structures. A flat consumer subscription does not make API calls free, and token pricing alone does not describe a coding-agent session. Cache use, tool calls, failed runs, retries and engineer review time determine the actual cost of a completed task.

For enterprise data, the model supports zero-data-retention arrangements like earlier Opus releases, but that does not replace a deployment review. U.S. organizations should still evaluate cloud region, logging, access control, regulated data, vendor terms and fallback behavior. Strong benchmark performance cannot decide whether a workflow is permitted to send its underlying data to a model.

Who should use it—and who probably should not

Opus 5.5 is most compelling for large codebase changes, multi-tool agents, long audits, professional research and complex financial or legal document workflows. Lower cache-read prices and stronger self-checking matter when a failed run creates expensive rework.

It can be excessive for short summaries, customer-support drafts, basic extraction and high-volume classification with deterministic validation. Output still costs $20 per million tokens, and maximum effort can consume very large reasoning and answer budgets.

The practical strategy is routing, not wholesale replacement. Begin with medium effort, reserve higher levels for workloads where they measurably lower failure and review rates, and use faster models for routine volume. Compare the cost of a finished result: API charges, retries, human review and latency together.

Bottom line

Claude Opus 5.5 is less a routine benchmark bump than an efficiency redesign of Anthropic’s top practical work model. Vendor results show large gains in agentic coding and knowledge work, while independent testing places it at the top of the launch-day Intelligence Index. The same independent data shows that maximum effort can use many more output tokens, and GPT-6 Astra still leads selected automation and science evaluations.

The useful question is not whether Opus 5.5 is universally “best.” It is whether a particular workload reaches the required quality at a lower total completion cost. For most teams, medium effort should be the baseline, with high, xhigh or max reserved for tasks where the reduction in failures is worth the extra time and tokens.

Sources and rights notice

Anthropic, Claude, OpenAI and GPT names and marks belong to their respective owners. This independent editorial article is not sponsored, endorsed or approved by either company. It separates vendor-reported figures from Artificial Analysis measurements and does not use official charts, product screens or third-party photographs.

Source: Anthropic · Includes original screenshots or graphics