Claude Opus 5.5: what changes for teams already running agents
In 60 seconds: Claude Opus 5.5 is Anthropic’s September 22, 2026 release for long-running agentic coding and knowledge work. Compared with Opus 5, the published API rate falls from $5/$25 to $4/$20 per million input/output tokens, cache reads fall from $0.50 to $0.20, and Anthropic reports more than 30% faster output. The migration is not a blind model-ID swap: thinking can no longer be disabled, forced tool choice returns an error, thinking blocks must be preserved, and some computer-use integrations need a new toolset. Run your own evals before moving production traffic.
Anthropic calls Claude Opus 5.5 the first model in the Claude 5.5 family. For a team that already has agents in production, the launch is less about another benchmark table and more about a changed operating envelope: lower published rates, different request constraints, and a few failure modes that can stop an existing agent before it calls its first tool.
These field notes use Anthropic’s announcement and platform documentation checked on September 30, 2026. Performance, cost reduction, and safety results below are provider-reported. They are starting points for an evaluation, not estimates of ROI for your workflow.
What changes from Opus 5
Lower rates and faster output, with an important caveat
Anthropic publishes these rates:
| API usage | Claude Opus 5.5 | Claude Opus 5 | Change |
|---|---|---|---|
| Input tokens | $4 / MTok | $5 / MTok | 20% lower |
| Output tokens | $20 / MTok | $25 / MTok | 20% lower |
| 5-minute cache writes | $5 / MTok | $6.25 / MTok | 20% lower |
| Cache reads | $0.20 / MTok | $0.50 / MTok | 60% lower |
Anthropic says a typical workload costs about 40% less at default settings because Opus 5.5 also uses fewer tokens per task, and that output generation is more than 30% faster. Treat both as Anthropic’s measurements, not a budget forecast. An agent’s bill still depends on effort, loop length, cache-hit rate, tool results sent back into context, retries, and how often a person asks it to redo work.
Replay the same representative cases and compare total cost per completed case. Price per token alone can hide a longer loop; a strong cache-hit rate can make the lower cache-read price matter more than the headline input rate.
The model now defaults to medium effort
Opus 5.5 uses adaptive thinking on every request, with medium as its default effort. Opus 5 defaulted to high. If your production configuration does not set effort explicitly, switching models changes two variables at once: model generation and reasoning depth.
Set the effort level in the test configuration instead of inheriting the default. Measure whether low or medium completes the task to your acceptance criteria, then reserve higher levels for cases that benefit from them. Thinking tokens count as output tokens, including when their text is not shown, and max_tokens covers thinking plus the final response.
Four breaking changes to review before the rollout
Anthropic’s what’s new guide lists four changes that can break code already running on Opus 5.
1. Thinking cannot be disabled
Both thinking: {"type": "disabled"} and the older manual budget form return a 400 error. Omit the field or use adaptive thinking, then control depth with output_config.effort.
Every response may begin with thinking blocks. Code that assumes content[0] is text will eventually fail. Select blocks by type, and in a tool loop return the assistant’s thinking blocks complete, unchanged, and in the original order.
2. Forced tool choice returns an error
tool_choice values any and tool are not supported. Use auto or none. If a downstream system requires schema-valid arguments, pair auto with strict tool use or use structured outputs; if the workflow requires a tool, say when it applies in the prompt and validate the result in application code.
This deserves an integration test. A workflow that previously forced create_invoice or update_crm may now answer in text unless the tool contract and prompt are clear. Do not convert that ambiguity into broader write permissions.
3. Thinking blocks belong to a model and conversation
Do not rebuild or edit old assistant turns while preserving only selected thinking blocks. The API verifies their signatures and, for some accounts, whether earlier messages, the system prompt, or tools changed before a replayed block. An edited, reordered, or partially dropped block can return a 400 error.
For routers, keep a clean conversation boundary when switching models. A fallback should not assume every model can consume the same preserved-thinking history.
4. The older computer-use tool is rejected on two platforms
On the Claude API and Google Cloud, Opus 5.5 rejects computer_20251124. Migrate to computer_toolset_20260801, remove the old beta header, and update the loop for the toolset’s result shape. Amazon Bedrock continues to accept computer_20251124, so this check is platform-specific.
Specs worth putting in the migration ticket
The current model overview lists:
| Item | Published value as of September 30, 2026 |
|---|---|
| Official name | Claude Opus 5.5 |
| Claude API model ID | claude-opus-5-5 |
| Context window | 1M tokens |
| Maximum output | 128K tokens |
| Thinking | Adaptive, always on |
| Default effort | medium |
| Input / output price | $4 / $20 per MTok |
| Cache read price | $0.20 per MTok |
| Availability | Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, Claude Platform on AWS |
| General latency SLA | Not published on the model page; verify your platform or contract |
| Account rate limits | Not a single model-wide value; verify the limit shown for your account and provider |
Fast mode is a separate research preview on the Claude API, priced at $8/$40 per million input/output tokens. It is not available on Bedrock, Claude Platform on AWS, Google Cloud, or Microsoft Foundry. Do not use its advertised speed to plan a standard-mode migration.
Safeguards and the agent’s permission model
Anthropic reports stronger automated behavioral-audit results and better prompt-injection resistance than Opus 5. It also says Opus 5.5 is less likely to take hard-to-reverse actions or go beyond the boundaries it receives. Those are useful provider results, but they do not replace scoped credentials, approvals, or an audit log.
There is also a concrete integration change: Opus 5.5 uses cybersecurity and biology safety classifiers. A declined request returns HTTP 200 with stop_reason: "refusal" and a stop_details category. If your loop treats every 200 as a successful answer, it may record an empty or incomplete task as done. Handle the stop reason explicitly, choose a safe fallback policy, and keep irreversible actions behind approval.
For a broader permissions pattern, use the copilot-or-autopilot framework: reading, recommendations, drafts, approved execution, and bounded autonomy should have different controls.
A migration checklist for production agents
- Inventory the route. Record the current model ID, cloud provider, thinking configuration, effort, tool-choice mode, computer-use version, beta headers, and fallback models.
- Create a separate test configuration. Change the platform-specific model ID without altering production traffic. Keep a rollback route to Opus 5 while you evaluate.
- Remove rejected request settings. Drop disabled/manual thinking, forced tool choice, non-default sampling parameters, and assistant prefills. Set effort explicitly.
- Fix response parsing. Read content blocks by
type; preserve every thinking block exactly in tool loops; decide whether user-facing progress needsthinking.displayconfigured. - Retest every tool contract. Use
autowith strict schemas or structured outputs where appropriate. Test the text-response path as well as successful, invalid, and denied tool calls. - Update computer use by platform. Move Claude API and Google Cloud integrations to
computer_toolset_20260801; do not copy that change blindly to Bedrock. - Handle refusals and limits. Branch on
stop_reason, inspectstop_details, budget enoughmax_tokensfor thinking plus text, and make fallback behavior observable. - Replay representative cases. Compare task completion, human correction, tool errors, refusals, end-to-end latency, total tokens, cache hits, and cost per accepted result. Keep the prompts and grading rubric fixed.
- Canary the rollout. Start with bounded traffic and read-only or approval-gated actions. Expand only after the error budget and review sample meet your team’s thresholds.
Migration time depends on the integration. An agent with one read-only tool and typed content parsing is a different job from an agent that edits conversation history, forces tools, and controls a browser.
What this means for a LATAM team
The lower USD rates matter, but exchange-rate exposure remains. Keep budgets in both tokens and local-currency reporting, and alert on cost per completed workflow rather than on raw token volume alone.
The cloud platform affects data location, support terms, quotas, and invoicing. Anthropic’s public model pages do not publish one regional latency SLA or one residency promise that covers every LATAM deployment. Verify those fields with the provider you actually use.
Test Spanish and mixed-language records from your own operation. A general multilingual capability does not prove that the model will preserve your product names, local legal vocabulary, address formats, or escalation rules.
The decision after the test
Opus 5.5 is a credible migration candidate when an Opus 5 agent is constrained by per-case cost, long loops, or output latency. The published pricing improvement is clear. The production decision still belongs to your eval set.
Keep the comparison narrow: same workflow, same tools, same permissions, explicit effort, and the same acceptance rubric. A separate field note will compare Claude Opus 5.5 with GPT-6.1 Sol; this one is intentionally about the Opus 5 to 5.5 migration.
If the agent itself is stable but its operating cost is hard to explain, first map the surrounding spend with the guide to the real cost of an AI agent. And if you are still on the previous generation, the Claude Opus 5 guide provides the earlier baseline.
Frequently asked questions
What is the Claude API model ID for Opus 5.5?
The fixed Claude API model ID is claude-opus-5-5, with no date suffix. Amazon Bedrock uses anthropic.claude-opus-5-5; Google Cloud, Microsoft Foundry, and Claude Platform on AWS use their documented platform identifiers.
Can a production agent disable thinking in Claude Opus 5.5?
No. Adaptive thinking is always on. Requests that disable thinking or set a manual thinking-token budget return a 400 error. Control depth, latency, and cost with effort, and test an explicit level because Opus 5.5 defaults to medium while Opus 5 defaulted to high.
Is migrating from Claude Opus 5 to Opus 5.5 only a model-name change?
Not for most tool-using agents. Teams must check thinking-block handling, forced tool choice, preserved conversation state, computer-use tool versions, refusal handling, token limits, and platform-specific model IDs before moving production traffic.
From insight to action
Want to turn this into an agent that works for your team?
Tell us which process you want to improve. In a free call, we will identify the first workflow worth building.