AI

Claude Opus 5.5 Makes Effort Tuning an Infrastructure Decision

Anthropic’s Claude Opus 5.5 arrives with the sort of headline that can make an infrastructure team reach for a model alias immediately: the company says it performs at the level of its higher-tier Claude Fable 5.1 on most work, while typical workloads cost 40% less than Opus 5. But the more consequential change is not a simple price cut. Opus 5.5 makes reasoning effort a first-class deployment control, keeps adaptive thinking on for every request, and changes several API assumptions that older agent loops may have baked into production.

That combination creates a clear operational lesson. Teams should treat Opus 5.5 as a new serving profile, not a drop-in model replacement. Its economics can materially improve long-running coding and knowledge-work agents, but only if operators re-baseline cost, latency, token ceilings, caching, and tool-loop handling at each effort level. Changing the model ID without doing that work risks replacing an expensive predictable workload with a cheaper-looking but poorly bounded one.

The efficiency claim is bigger than a token-price cut

Anthropic prices Opus 5.5 at $4 per million input tokens and $20 per million output tokens, down from $5 and $25 for Opus 5. Cache reads fall from $0.50 to $0.20 per million tokens, while five-minute cache writes fall from $6.25 to $5. Those list-price changes matter, especially for agents that repeatedly reuse a large system prompt, repository map, policy corpus, or conversation history.

Yet Anthropic’s 40% typical-workload savings claim is larger than the 20% reduction in base input and output prices. The company attributes the difference to Opus 5.5 using fewer tokens per task as well as charging less per token. That distinction should shape how platform teams validate the model. A price-sheet comparison estimates the cost of the same token stream; an agent evaluation measures whether the new model takes fewer steps, issues fewer tool calls, retries less often, and reaches an acceptable result sooner.

Anthropic says Opus 5.5 generates output more than 30% faster than Opus 5. It also offers a fast mode with up to 2.5 times the speed, priced at $8 per million input tokens and $40 per million output tokens. Fast mode therefore is not simply “better latency” to enable globally. It is a separate service tier whose value depends on whether reduced wall-clock time improves an interactive workflow, increases completed jobs per hour, or prevents timeouts enough to justify roughly double the standard token rate.

The company’s own examples illustrate why task-level measurement is essential. Anthropic reports that an early tester’s 200,000-line codebase audit finished in under three hours, compared with more than 20 hours on Opus 5, while using 2.5 times fewer tokens. In an internal C-to-Rust translation of HAProxy, Opus 5.5 reportedly finished in 9.5 hours versus 12 hours for Fable 5.1 and cost 51% less. These are vendor-reported results, not a guarantee for another codebase, but they identify the right unit of analysis: cost and time per accepted task.

Effort is now part of the routing policy

Opus 5.5 supports five effort levels—low, medium, high, xhigh, and max—and defaults to medium. Adaptive thinking is always on. Operators cannot disable it or specify a manual thinking-token budget. That design shifts control from a fixed reasoning budget to a policy decision: how much effort should a request receive, given its business value, complexity, latency objective, and retry cost?

This is an infrastructure concern because the same model can occupy several positions in a routing ladder. Low or medium effort may suit routine code review, document synthesis, and common support workflows. High effort can be reserved for tasks that fail a first pass or carry a larger cost of error. Xhigh and max may be appropriate for bounded, high-value jobs such as migrations or difficult incident analysis, provided queues, timeouts, and output ceilings can accommodate them.

A practical routing policy should use evidence rather than prompt adjectives. Record the effort setting with every trace, then compare it against:

  • Quality: pass rate on workload-specific evaluations, human acceptance, and downstream corrections.
  • Task cost: uncached input, cache writes, cache reads, output tokens, tool costs, and retries.
  • Latency: time to first useful update, time between tool calls, and end-to-end completion time.
  • Agent behavior: number of steps, repeated tool calls, loop termination, and escalation frequency.
  • Reliability: refusal rate, malformed tool arguments, context exhaustion, and timeout rate.

The important comparison is not “medium versus max” in the abstract. It is the lowest effort setting that clears the service’s quality bar, plus a controlled escalation path. If low effort solves 90% of a task class and high effort recovers most of the rest, a two-stage router may be cheaper than sending every job at high effort. Conversely, retries can erase the savings if the first attempt consumes a long context or triggers costly external tools.

The API migration has real breaking points

Opus 5.5 uses the fixed model ID claude-opus-5-5 on the Claude API, with platform-specific IDs for Amazon Bedrock, Google Cloud, Microsoft Foundry, and Claude Platform on AWS. The model provides a one-million-token context window and up to 128,000 output tokens for synchronous Messages API calls. Those generous limits do not remove the need for guardrails: the request’s max_tokens covers both thinking and visible text.

That last detail can surprise applications migrating from configurations where thinking was disabled. Because adaptive thinking always runs, a token ceiling sized only for the visible response may truncate complex work. At xhigh or max effort, Anthropic recommends starting with at least 64,000 maximum tokens and tuning from observed behavior. Thinking tokens are billed as output tokens even when the thinking text is not shown, so dashboards that count only rendered text will understate usage.

Response parsing also needs inspection. An Opus 5.5 response can begin with one or more thinking blocks before the first text block. Code that assumes content[0].text will fail. Stream consumers should branch on block type, and tool-use loops must return thinking blocks unmodified with tool results. Filtering, reordering, or reconstructing those blocks can cause the API to reject the request. By default, thinking display text is omitted, which may look like a long pause in an interactive product; teams that need visible progress can request summarized thinking.

Other request settings are deliberately constrained. Forced tool selection using tool_choice values such as any or a named tool is rejected; applications should use automatic tool selection and enforce output requirements through strict tools or structured outputs. Non-default temperature, top-p, and top-k values are rejected. Prefilled assistant turns are also rejected. Computer-use integrations on the Claude API and Google Cloud must move to the newer toolset declaration documented for the model.

These changes are not incidental SDK cleanup. They affect gateways, observability middleware, tool executors, replay systems, and fallback routers. A gateway that normalizes every provider response into a text-first schema may silently discard the state needed for the next tool call. A fallback system can also move from a model that always emits thinking blocks to one that does not, so conversation state and parsing must tolerate both shapes.

Caching becomes the economic center of agent workloads

The steepest unit-price reduction is on cache reads: 60% below Opus 5. For agents that repeatedly carry a stable prefix, cache reads may dominate the input side of the bill. Anthropic lists Opus 5.5 cache reads at 5% of the base input-token price, compared with the more common 10% multiplier used by many other Claude models.

That makes prompt topology an infrastructure optimization. Stable instructions, tool definitions, repository summaries, and reference material should be arranged before volatile conversation content so they can be reused. But a low cache-read price does not justify caching everything. Five-minute writes cost $5 per million tokens and one-hour writes cost $8, so the break-even point depends on reuse frequency and cache lifetime.

Teams should expose cache creation and cache-read tokens separately in cost telemetry. A blended “input tokens” metric hides whether an agent is benefiting from reuse or constantly invalidating its prefix. For bursty jobs, batch scheduling may create enough temporal locality to improve cache hit rates. For sporadic jobs, aggressive one-hour cache writes may cost more than they save.

A migration plan for production teams

1. Freeze the current baseline

Before changing the model, capture accepted-task rate, p50 and p95 completion time, tokens by category, tool-call count, retry rate, and total cost per successful task for Opus 5. Include failure classes, not just averages. Long-running agents often have a small tail of runaway jobs that determines operational cost.

2. Run an effort sweep on representative work

Replay a fixed evaluation set at low, medium, high, xhigh, and max effort where appropriate. Use the same tool sandbox and external-service limits as production. Judge final artifacts, not only benchmark-style answers. The output should be a routing table that names a default effort and escalation rule for each workload class.

3. Validate the agent loop before quality testing

Update parsers to select content by block type, preserve thinking blocks across tool calls, handle refusals, and remove rejected request fields. Test streaming, cancellation, timeouts, fallback transitions, and replay. This separates integration failures from model-quality failures and prevents a broken harness from contaminating the evaluation.

4. Set layered budgets

Use a per-request token ceiling, a per-task spending limit, maximum tool calls, wall-clock timeout, and loop-detection rule. A one-million-token context is capacity, not a target. Summarize or retrieve older context when doing so preserves quality, and reserve very large contexts for tasks that demonstrably benefit.

5. Roll out with shadow traffic and a rollback path

Start with non-mutating workloads or shadow evaluations. Then canary a narrow slice of production traffic while comparing task-level economics. Pin the model ID rather than assuming an alias will preserve behavior, and keep the previous route available until the new model meets both quality and budget objectives over the workload’s normal variance.

What changes for AI infrastructure

Opus 5.5 reinforces a broader change in model serving: the model name is no longer a sufficient routing key. Effort, cache state, latency mode, context size, tool access, and risk tier jointly define the actual service profile. Two requests sent to the same model can have very different cost and latency characteristics because one runs at low effort over a reusable prefix while another runs at max effort through a long sequence of tools.

For platform teams, the opportunity is meaningful. Lower list prices, cheaper cache reads, and fewer reported tokens per completed task can make frontier-level agents viable for work that previously failed a cost threshold. The obligation is equally clear: measure the complete task, preserve the protocol state, and make effort escalation explicit. Opus 5.5’s strongest infrastructure feature is not that it is cheaper by default; it is that teams now have a more legible control surface for buying additional reasoning only where it earns its keep.

Sources