Field report

A cheap run is not a cheap task

A model with a low token price can still be the expensive choice. What matters is how many runs, agent steps, reviews and corrections it takes before the result is actually accepted.

In my first Agentic Coding Digest I wrote:

Smarter → needs less reasoning → fewer tokens for easy work → cheaper.

That is directionally right, but to be honest, it’s oversimplified.

I was still looking at the price of one run. But a cheap run and a cheap task are not the same thing. A run can cost almost nothing and still be the most expensive path to a result if it fails, needs three follow-ups and leaves you with a diff you have to repair yourself.

So the short version now is:

Don’t optimize for the cheapest run. Optimize for the cheapest accepted result.

TL;DR

  • Token price is the wrong unit → what matters is what the complete task costs
  • Capability first, effort second, token price third
  • Effort helps a model use the capability it has, it does not create capability
  • A cheap failed run is only the first installment
  • Put capability where the uncertainty is, then let cheaper models execute the decided parts

Capability first, effort second

There are two controls we usually look at: model tier and effort. The model decides what kind of problem it can solve. Effort decides how much work it invests into solving it.

And “work” here is more than the invisible thinking before an answer. It is understanding the request, exploring the codebase, trying tools, testing assumptions, reviewing the change and fixing what it finds before it comes back to you.

More effort can improve the result. It can also mean more exploration, more tool calls, more things you never asked for and a lot more tokens.

The important boundary is:

More effort does not equal more capability.

Effort helps a model to use the capability it already has. It does not magically give it capabilities it doesn’t have.

That’s why “take the cheapest model and raise the effort until it works” can backfire. The model may burn more steps trying to compensate and still not reach the judgment of a stronger model at lower effort.

But the opposite rule, “always take the strongest model”, is not better. A fully specified migration does not become more correct because the most capable model renamed the fields. You already made the decisions. There is nothing left to spend judgment on.

A cheap failed run was not the cheap solution

The price shown in a model picker tells you what one run costs. Your task normally does not end there.

Maybe the first result misses the point. You read it, explain the problem and run it again. Then another model reviews it. The review finds something, so the original agent goes back into the code. At some point you finally accept the result.

The real cost is closer to:

all runs until accepted + review + rework

A cheap failed run was not the cheap solution. It was only the first installment.

This is also why I don’t think token efficiency and model intelligence can be separated from behavior. Does the model stay on task? Can you steer it during a run? Does it follow a changed direction without forgetting the work already in progress? Does its way of writing code match what you expect?

A model can have a low token price and good benchmark scores and still be expensive for you because you don’t trust what it does between the prompt and the result.

Put capability where the uncertainty is

So does that mean the strongest model should do every part of every difficult task? No. That gets expensive in the other direction.

A task needs capability while it is still unclear: understanding what is actually wanted, choosing the approach, deciding what “done” means, splitting the work and judging the results.

Once these decisions are made, the leftover work looks completely different to a worker model. It is small, bounded and verifiable. The task complexity is not fixed. A capable model can shrink it before a cheaper model ever touches the task.

That’s the part I missed in my first explanation. The strong model does not only solve problems, it removes uncertainty from tasks. And a task without uncertainty needs way less capability.

The split I use now looks like this:

  • Capable model: clarify, plan, split, write the worker briefs and review what comes back
  • Cheaper model: search, implement and test the clearly defined parts
  • Worker starts struggling: don’t keep raising the effort and hoping. Re-plan the task or escalate the model

That is also the economic reason behind putting a strong model in the orchestrator seat. It spends its expensive tokens changing the shape of the work, not typing every decided line itself.

My current rule of thumb

  • Trivial, deterministic work: cheap model, low effort
  • Non-trivial work: a model clearly capable enough, then the lowest effort that solves it reliably
  • Ambiguous or architectural work: spend capability on the plan and the decisions, then hand the bounded work down
  • Repeated failures: re-check the model-task fit instead of blindly rerunning the same configuration
  • Maximum effort: use it intentionally, not because it feels safer

These are still rules of thumb. The task, model, harness and your own setup all change the result. That is exactly why the accepted result is the useful unit. It includes the complete system instead of pretending the model dropdown is the whole story.

I already had to correct my own example

There is a nice little problem with writing model advice publicly: the example can become wrong before the mechanism does.

In ACD #2 I wrote that Opus 5 felt like a good default for non-trivial work. By ACD #3 I had stopped selecting it. The weird behavior I mentioned in the first take became more common, and the direction became hard to trust.

I could quietly remove the old recommendation and replace it with another model name. But then this article would make the same mistake again. The point was never that one specific expensive model always wins. The point is that price per token tells you very little about how much work sits between your prompt and something you are willing to accept.

Judge the output, not the price tag. And count the whole task.

@edhorEnd of report

Back to all blog posts