Opus 5.5 is 2× the price on paper
Every column on the price sheet says Opus 5.5 costs twice what Sonnet 5.5 costs. Every column but one — and in a long agent session that one column is most of your bill. Here is where the 2× actually lands, and the switching cost nobody prices in.
The price sheet is unambiguous. Claude Opus 5.5 is $4 per million input tokens and $20 per million output. Claude Sonnet 5.5 is $2 and $10. Twice the price, in both directions, no asterisk. Which is why every tier conversation I have been in this month has ended at the same place: use Sonnet, escalate to Opus when it matters, said with the confidence of someone who has read one table.
A piece by @gippp69 on X nudged me into actually pricing a session rather than a token, and the number that falls out is not 2×. In a long agent session it is closer to 1.2×. The reason is one row of the price sheet that does not double, and it is the row that most of your tokens go through.
The row that does not double
| Per MTok | Opus 5.5 | Sonnet 5.5 | Gap |
|---|---|---|---|
| Base input | $4.00 | $2.00 | 2× |
| 5-minute cache write | $5.00 | $2.50 | 2× |
| 1-hour cache write | $8.00 | $4.00 | 2× |
| Cache read | $0.20 | $0.20 | 1× |
| Output | $20.00 | $10.00 | 2× |
Cache reads are normally 0.1× a model's base input price. On Opus 5.5 the multiplier is 0.05× — half the usual discount rate, which drops its cache reads to exactly what Sonnet 5.5 charges for the same thing. Whatever the reasoning behind that, the consequence is concrete: reading context costs the same on both models.
Now think about what a turn deep inside an agent session actually consists of. A few thousand fresh tokens — the new user message, the tool results since the last breakpoint. A couple of thousand output tokens. And three hundred thousand tokens of cached prefix, replayed, because the API is stateless and the whole conversation goes back up the wire every single turn. The thing that dominates the token count is the thing that is priced identically.
Four turns, priced
Same shape each time: a cached prefix read back, 4,000 fresh input tokens, and an output. Only the depth of the session and the length of the answer change.
- Opus 5.5
- Sonnet 5.5
Turn 2 — 20K cached, 2K out
effective gap 1.88×
Turn 40 — 300K cached, 2K out
1.32×
Turn 90 — 600K cached, 2K out
1.19×
300K cached, 8K out
1.59× — output pulls it back
The sticker gap is 2× and you only pay it at the start. By turn forty the premium for the better model is about three pence a turn. Note the last bar: the gap is a function of how much your turns write, not how long they are.
Read the trend rather than the exact figures, because your prefix and your output lengths are not mine. The structure holds regardless: as a session deepens, the Opus premium decays toward 1×, and the only thing holding it up is output tokens. A session that reads enormously and writes briefly — code review, search, triage, most tool-heavy agent work — converges fast. A session that generates whole files on every turn stays near the sticker price.
Which also means the standard advice is backwards for the standard case. "Start cheap, escalate when it gets hard" assumes the premium is constant. It is largest exactly when the session is young and nothing has gone wrong yet, and smallest at turn ninety when you have finally decided the task is hard.
The switch is the expensive part
Here is the cost that never appears in these comparisons, because it is not a per-token rate. Prompt caches are scoped to a model. Sonnet's cache of your 300K-token session is not readable by Opus. Switch mid-session and the whole prefix is written again from scratch, at the new model's write price.
- to re-cache a 300K session on Opus 5.5$2.40 on the 1-hour TTL
- $1.50
- turns of Opus premium that buysdepending on output length
- 17–54
- cache-read gap between the two models$0.20 on both
- 1×
- cheaper than Opus 5 on a cache-heavy turnnot the headline 20%
- 47%
That $1.50 is 300,000 tokens at the $5 five-minute write rate. Measured against the per-turn premium from the chart — about 2.8¢ on a light turn, 8.8¢ on a writing-heavy one — a single mid-session switch costs somewhere between seventeen and fifty-four turns' worth of the upgrade you were trying to avoid paying for. You did not save money by starting on Sonnet. You spent the savings, plus interest, on the moment you changed your mind.
A new session starts
The cache is empty. Switching now costs nothing.
Is this long-horizon, high-stakes, or hard to verify?
↳ yes? start on Opus and stay there. The premium shrinks every turn; the switch never does.
Decide within the first handful of turns
While the prefix is small enough that re-caching is pocket change.
Prefix passes ~100K tokens
↳ the window has closed — switching now costs more than finishing on whatever you are on.
Changed your mind anyway?
Hand off a clean brief to a fresh session. Do not replay the transcript into the new model.
One model, one cache, for the life of the session
The handoff at the end is the useful trick. A 2,000-token summary written into a new Opus session costs about a penny to cache; dragging 300K tokens of Sonnet transcript across costs $1.50 and brings all the dead ends with it.
Two invalidators that cost more than the tier
Having established that re-caching is the expensive event, the obvious follow-up is: what else triggers one?
- Changing top-level
effortmid-conversation. It invalidates the messages cache — same $1.50, no model change required. On Opus 5.5 and Sonnet 5.5 there is a per-message effort system message in beta (an empty-content{role: "system"}entry carryingoutput_config, behind themid-conversation-output-config-2026-07-01flag) that changes effort from that point on without resetting the prefix. Use it. - Opus 5.5's effort default is
medium, where Opus 5's washigh. Code migrated from Opus 5 that never seteffortexplicitly is now running a level lower. That is a quality change and a spend change arriving together, silently, in the column that drives the whole 2× gap. Set it explicitly and stop guessing which one you are on.
// The price sheet, $ per MTok. Only one row is not 2x.
const PRICE = {
"claude-opus-5-5": { input: 4, cacheWrite: 5.00, cacheRead: 0.2, output: 20 },
"claude-sonnet-5-5": { input: 2, cacheWrite: 2.50, cacheRead: 0.2, output: 10 },
} as const;
// Your real ratio is in the usage block the API already hands back.
// input_tokens excludes anything served from cache — don't double-count it.
function turnCost(model: keyof typeof PRICE, u: Anthropic.Usage) {
const p = PRICE[model];
return (
u.input_tokens * p.input +
(u.cache_creation_input_tokens ?? 0) * p.cacheWrite +
(u.cache_read_input_tokens ?? 0) * p.cacheRead +
u.output_tokens * p.output
) / 1_000_000;
}
// Log it per turn. The number you want is dollars per passing run,
// not dollars per turn — a cheaper turn that needs three more is not cheaper.Two things the headline numbers hide
The 20% cut was much bigger than 20%. Opus 5 was $5/$25; Opus 5.5 is $4/$20 — that is the announced 20%. But the cache read went from $0.50 to $0.20, because the multiplier moved from 0.1× to 0.05×. On the cache-heavy turn above, Opus 5 cost 22¢ and Opus 5.5 costs 11.6¢. That is 47% cheaper, not 20%, and the gap widens the longer your sessions run. If you benchmarked Opus against Sonnet before this release and filed the result, the file is wrong.
Batched Opus 5.5 costs exactly what unbatched Sonnet 5.5 costs. The Batch API halves both directions: $2 in, $10 out — the same two numbers as Sonnet 5.5 at full price. For anything that tolerates asynchronous processing — overnight evals, bulk extraction, backfills, regression suites — the tier question simply dissolves. You were never choosing between models there. You were choosing between a queue and a worse answer.
What to actually do
- 1Price a session, not a token. Pull
usagefrom your own logs and compute cost per turn at the depth your sessions actually reach. The ratio you care about is yours, and it is not 2×. - 2Decide the model in the first few turns, while the prefix is small and the decision is nearly free.
- 3Never switch mid-session. If you must change your mind, write a clean brief into a fresh session instead of dragging the transcript across.
- 4Set
effortexplicitly, and change it with a per-message system message rather than the top-level field. - 5Batch anything that can wait. Opus at Sonnet's price is not a tradeoff.
- 6Judge on dollars per passing run. A turn that costs 25% less and needs two more turns to get there cost you more.
The sticker price compares tokens. Your invoice compares sessions. They are not the same comparison.
None of this makes Opus 5.5 the right default for everything — plenty of work is genuinely Sonnet-shaped, and some is Haiku-shaped. What it kills is the specific reflex of reaching for the cheaper model because a table said 2×, in a workload where the real figure is 1.2× and the act of changing your mind later costs more than the entire difference.
Price one real session before you tier anything. The arithmetic takes ten minutes and it is the only version of this comparison that is about your application rather than someone's pricing page.
Building something like this?
I design and ship these systems for clients: retrieval over private data, agents that complete real tasks, and the Laravel platforms underneath them.