For two years the standard way to cut inference bills was to pick a cheaper model. Route the easy stuff down, cache what repeats, and keep the frontier model for the hard remainder. That playbook still works. It just stopped being where the money is.
Anthropic spent three releases in three weeks moving the price surface inside the model. Opus 5.5 landed September 22 with five effort levels and cost-per-attempt charts for each setting. Sonnet 5.5 followed on September 28, plotting every benchmark at cost per task, with the framing stated outright: as effort goes up, models typically work for longer, leading to a higher cost per task but generally also a higher score. Haiku 5.5 arrived October 7 as the first Haiku with the adjustable dial. Three models, three weeks, one pattern. The tier that matters is no longer which model you called. It is how hard you asked it to think.
The Sonnet numbers make the point concrete, and they are Anthropic’s own. Anthropic says Sonnet 5.5 at Low and Medium effort beats the best Sonnet 5 score for roughly a tenth of the cost per task. Read that again. Not a smaller model beating a bigger one. The same model, told to think less, beating the previous generation at a tenth of the price. Anthropic also advises running lower effort for routine work, which is the vendor telling you the dial is the product.
I trust the shape of that curve more than any point on it, because the ends of the dial are wilder than a chart suggests. Simon Willison tested both ends independently. His Haiku 5.5 runs put a low-effort SVG at 7 seconds and about a tenth of a cent against a max-effort run at over five minutes and 3.38 cents. Same model, same prompt shape, roughly a fortyfold cost gap. And his Opus 5.5 test found max effort spending 128,000 thinking tokens on a single SVG, then failing from overthinking. Thinking harder can make the answer worse while costing more. That failure mode does not exist in a per-token price list.
So here is my claim. Effort is a cost dial you set per tool call, not per deployment, and almost nobody does. Every agent run I would build is a mix of trivial calls and genuinely hard ones. Classifying a request, picking a tool, checking a format: these want the floor setting. The one planning step the whole run depends on wants the ceiling. A single global effort level prices every call as if it were the average call, and the average call does not exist.
The design I would build starts from a per-call effort budget. Each tool in the loop declares a default effort. Cheap gates sit at low. The planner and the final synthesis sit high. Two rules keep it honest. Retries escalate: a failed low-effort attempt goes again one notch up instead of starting high everywhere just in case. And the run carries a thinking budget, so a runaway step cannot spend the whole allowance before the steps after it get a turn. When the budget runs thin, remaining calls drop a notch rather than failing. This is guesswork about the best policy, not a measured result, but the mechanism is simple enough to hold: effort in, cost out, score usually following.
What I would check first is the distribution, not the average. Log thinking tokens per tool call for a week at a fixed medium setting. My guess is the familiar shape: a fat head of calls where low effort would have scored the same, and a thin tail where max effort earns its keep. If that shape shows up, the budget practically writes itself. Put the floor under the head and the ceiling over the tail, and the tenth-of-the-cost result Anthropic reports starts looking like something you can reproduce rather than a vendor chart.
Two years of routing and caching treated every model as a fixed price. The dial ends that. Price by thinking, and spend it where the thinking matters.
