Skip to content
Cost Engineering7 min read·

Not every call needs the expensive model

Tiering your agent's calls is the obvious cost lever and usually the second-best one. Here is the arithmetic on which calls earn the frontier model and which never did, priced on both the Anthropic and OpenAI ladders — where caching beat routing on each of them.

Model RoutingCost OptimisationPrompt CachingAgent ArchitectureOpenAIGPT-6 AstraHaiku 4.5

Every agent I have put into production has had the same cost curve. Negligible, negligible, negligible, then a Monday where someone forwards you the invoice with nothing written in the body of the email. The reflex at that point is to swap the cheap model in globally and see what breaks. The other reflex — leave everything on the frontier model, because that is what the demo ran on — is the same mistake with a different invoice.

Both share an assumption worth dismantling: that a run has a model. It does not. A single agent run is twenty calls, or two hundred, and they are not the same kind of work. One decides the plan. Three write something a human will actually read. The other sixteen are formatting tool arguments, deciding whether a file is relevant, pulling four fields out of a PDF, or summarising a diff nobody will ever open. Paying frontier prices across all twenty because four of them needed it is the bug.

Price out a real run before you touch anything

Abstract advice about tiering is cheap. Arithmetic is not. So here is a shape I keep meeting — one run of a code-review-ish agent, twenty model calls, two populations:

  • Four heavy calls. One planning pass (20K in, 3K out) and three synthesis passes (40K in, 2K out each). Long context, long output, consequential.
  • Sixteen light calls. Classification, relevance filtering, schema extraction, tool-argument formatting (6K in, 400 out each). Narrow, mechanical, checkable.

That is 236,000 input tokens and 15,400 output tokens per run. Note what it also is: the four heavy calls are 20% of the call count and 59% of the tokens. This is the thing that wrecks most tiering plans before they start. The calls you are itching to downgrade are the many, but the tokens live with the few.

I will price it twice, because the answer turns out not to depend on whose models you buy — and the two ladders are priced very differently indeed.

TierAnthropicOpenAI
Frontierclaude-opus-5-5 — $4 / $20gpt-6-astra — $10 / $50
Midclaude-sonnet-5-5 — $2 / $10gpt-6.1-sol — $2 / $10
Cheapclaude-haiku-4-5 — $1 / $5gpt-6-luna — $0.10 / $0.50
Frontier cache read$0.20, a 20× discount$1.00, a 10× discount
Frontier-to-cheap gap4×100×

Hold on to the last two rows. Anthropic's ladder has a shallow tier gap and a steep cache discount; OpenAI's has the reverse, on both counts. If anything was going to make routing the obvious first move, it is a hundred-to-one price gap.

of the calls in this runbut 59% of its tokens
20%
saved by routing alone40% on the OpenAI ladder
31%
saved by caching alone54% on the OpenAI ladder
57%
OpenAI frontier-to-cheap gapagainst 4× on the Anthropic ladder
100×
Chart 1: the same workload, as a share of each ladder's all-frontier bill
  • Anthropic
  • OpenAI

Frontier model throughout

the demo configuration

Anthropic100%
OpenAI100%

Tiered, no caching

16 of 20 calls on the cheap tier

Anthropic69%
OpenAI60%

Frontier throughout, cached

no routing at all

Anthropic43%
OpenAI46%

Tiered and cached

both levers

Anthropic35%
OpenAI27%

Cheap tier throughout

cheapest, and it does not finish the job

Anthropic25%
OpenAI1%

Straight arithmetic on published rates — Opus 5.5 at $4/$20 per MTok against Haiku 4.5 at $1/$5, and GPT-6 Astra at $10/$50 against GPT-6 Luna at $0.10/$0.50 — assuming 80% of the heavy calls' input is a stable prefix, read at $0.20 and $1.00 respectively. Shown as a share of each ladder's own all-frontier bill, because the shape is the point. Cache writes are excluded on the grounds that they amortise over repeat reads; if your prefix changes every run they do not, and caching is not your lever.

Read the middle two bars together, because they are the whole post, and because they say the same thing on both ladders. Moving sixteen of twenty calls down a tier — the intervention everybody reaches for first, the one that needs a router, an eval per route and a new failure mode — saved 31% on the Anthropic ladder and 40% on the OpenAI one. Changing nothing about which model runs and simply caching the stable prefix saved 57% and 54%. In money: $1,252 per thousand runs down to $442 on one ladder, $3,130 down to $855 on the other. Routing is a real lever. It is just not the first one, and even across a hundredfold tier gap it was not the biggest.

Which calls actually earn it

Once the free wins are banked, the question becomes per-call rather than per-run. The dividing line is not difficulty. It is whether a mistake is visible, and what it costs you when it is not.

Call in the runTierWhy
Planning and decompositionFrontierAn error here multiplies through every later step, and nothing downstream will catch it
Final synthesis for a humanFrontierIt is the only output anybody reads. Degrade this and you have degraded the product
Arguments for an irreversible actionFrontierA write, a payment, a delete. The cost of error is unbounded
Subagent that reads 40 files and reports 10 linesCheapVolume in, little out. The judgement is in the question, not the reading
Classification, routing, triageCheapNarrow, checkable, and the thing small models are genuinely good at
Structured extraction to a schemaCheapA strict tool schema does the enforcing, not the model's good intentions
Relevance filter or rerankerCheapA miss costs one wasted retrieval, and the next step notices
Commit messages, titles, per-file notesCheapNobody is harmed by a mediocre commit message
A critic reviewing the frontier model's workFrontierA weaker critic rubber-stamps. Anthropic's advisor tool enforces this in the API — an advisor less capable than the executor is a 400
Flowchart 1: which tier does this call belong in?
  1. One call in the run

    Decide per call kind, never per run.

  2. Does a wrong answer corrupt later steps?

    ↳ yes? frontier. The plan and the irreversible actions live here, permanently.

  3. Can something other than a model check the output?

    ↳ no? frontier. With no check you will never notice the day it got worse.

  4. Is it a transform, a classification or a bounded extraction?

    Narrow, structured, repetitive.

    ↳ yes? cheapest tier that clears its own eval, with the check wired in.

  5. Everything else: capable model, lower effort

    Measured on a sample of real traffic, not chosen by vibe.

  6. A route table you can read in one screen

The second check is the one people skip, and it is the one that matters. A downgrade with no verification attached is not a cost saving. It is a quality change you have agreed not to measure.

What this looks like in code

Keep the routing declarative and in one place, and name the tiers rather than the models. The moment a model id is scattered across forty call sites, nobody can answer "what are we spending this on" without grepping, and changing vendor stops being a one-line decision.

ts
// Bind the ladder once. Swapping vendor is this block, not the next one.
const TIERS = {
  frontier: "claude-opus-5-5",   // or "gpt-6-astra"
  mid:      "claude-sonnet-5-5", // or "gpt-6.1-sol"
  cheap:    "claude-haiku-4-5",  // or "gpt-6-luna"
} as const;

// One table. Every routing decision in the system is visible here.
export const ROUTES = {
  // The few that justify the price.
  plan:      { tier: "frontier", effort: "high" },
  synthesis: { tier: "frontier", effort: "high" },
  mutate:    { tier: "frontier", effort: "high" },   // irreversible
  critique:  { tier: "frontier", effort: "medium" }, // never below the author

  // The many. Cheapest tier that clears its own eval.
  classify:  { tier: "cheap" },
  extract:   { tier: "cheap", strict: true },
  rerank:    { tier: "cheap" },
  readFiles: { tier: "mid", effort: "low" },         // volume, some judgement
} as const;

Then the escalation path, which is what makes a downgrade safe rather than merely cheap:

ts
async function routed(kind: Kind, input: Input) {
  const cheap = await call(ROUTES[kind], input);

  // The check is the entire design. No check, no downgrade.
  if (verify[kind](cheap)) return cheap;

  metrics.escalated(kind);           // the number that decides whether this pays
  return call(ROUTES.plan, input);   // one more attempt, on the expensive one
}

A cascade only pays while escalation stays rare, and how rare depends on which ladder you are standing on. With a cheap attempt costing c, an expensive one costing e, and an escalation rate p, you spend c + p·e against a flat e — so the cascade wins while p stays under 1 − c/e. The four-to-one gap between Opus 5.5 and Haiku 4.5 allows escalation up to 75%. The hundred-to-one gap between Astra and Luna allows 99%, at which point the arithmetic has stopped being the constraint and the only live question is whether you can verify the cheap answer at all. Neither figure prices the latency you now pay twice, or the wrong answers that quietly satisfy your check.

Escalation rate is not a metric you watch out of curiosity. It is the only thing standing between a cost optimisation and an elaborate way of paying twice.

Measure cost per completed task, not per call

This is where routing work usually goes wrong, and it is a measurement problem rather than an engineering one. A cheaper call that needs two more turns, or fails its check and escalates, or produces something a human sends back, is not cheaper. It is the same money, rearranged to look better on a per-token dashboard.

Log, per call: the route, the model, the effort, input and output tokens, cached tokens, whether it escalated, and whether its run finished. Then divide spend by completed tasks. Four numbers fall out, and they beat any amount of reasoning about which model is cleverer — cost per completed task, escalation rate per route, cache hit rate, and the share of spend by tier. Without them, every tiering decision is a guess delivered in a confident tone.

One more free win while you are in there. Anything that does not have to answer now — nightly summarisation, backfills, eval runs, document ingestion — belongs on the Batch API, which is half price at both vendors. No routing, no quality tradeoff, no new failure mode. A queue and some patience.

The order to do this in

  1. 1Measure first. Cost per completed task, and the share of tokens by call kind. Rarely where you assumed.
  2. 2Take the free wins. Cache the stable prefix, tidy the input tokens, move offline work to batch. No quality tradeoff, and on both ladders above a bigger win than routing.
  3. 3Then effort. Lower it per route, on the model you already run. One cache namespace, one eval, one field changed.
  4. 4Then route, starting with the narrow mechanical calls that have a real check attached. One at a time, watching the escalation rate.
  5. 5Leave the plan, the synthesis and the irreversible steps alone. They are 20% of your calls and most of your product.

The punchline

"Use a cheaper model" is not a strategy, and neither is "use the best model". The useful question is which of this run's calls are load-bearing, and the honest answer is usually about four of them.

Price out one real run. Cache the prefix. Drop the effort. Route the mechanical calls with a check wired in, leave the consequential ones where they are, and judge the result on cost per completed task. That sequence, in that order, is worth more than any amount of agonising over the model table.

Building something like this?

I design and ship these systems for clients: retrieval over private data, agents that complete real tasks, and the Laravel platforms underneath them.

Keep reading