Orchestrate, don't operate
Put capability where the uncertainty is. Let the strongest model clarify, split and judge the work, then hand the bounded parts to cheaper models.
The first time I got access to a model with a completely different level of taste, I used it exactly like the one before it. Same prompts, same single-agent loop, only much more expensive.
The model was not the problem. I put it in the wrong seat.
TL;DR
- Don’t run a new top model like you ran the old one → same usage, way more cost, similar output
- Its value is judgment: removing uncertainty before the worker starts
- Let it clarify, split, write the worker briefs and judge what comes back → let cheaper models implement and test the bounded parts
- If a worker struggles, re-plan or escalate instead of feeding it more and more effort
- If your tasks are already well defined, you are the orchestrator → skip the layer
The trap
The failure mode is sneaky because nothing looks wrong. You give the best model a task, it does the whole thing itself: greps the repo, reads twenty files, writes the boilerplate, runs the tests, fixes the lint… and the result is good.
It’s just that most of those steps produced output you can’t tell apart from what a model at a fraction of the price would have produced. And you paid the premium rate for all of them.
Writing the fifteenth similar handler is not a task that rewards taste. But exploration is not always interchangeable grunt work either. A capable model can save a lot of it by deciding which files matter, whether that fifteenth handler should exist at all, and what evidence would prove the result.
The waste is paying for judgment and then giving the model no role in the decisions.
Put it in the orchestrator seat
The strongest model you have holds the goal and clarifies what is still unclear. It writes the briefs for the agents doing the work, reads what comes back and decides what to accept, redo or throw away.
This is not a ban on the orchestrator touching the implementation. If the decision and the code are tightly connected, splitting them can be more expensive than doing them together. The point is to keep premium capability on the parts where its judgment changes the result.
Concretely that means telling it to:
- use cheaper models for helper tasks → searching a known area, mechanical implementation with a clear spec and repeatable tests
- start with the lowest effort that lets the model steer reliably → in my current setup that is usually medium. Treat that as a dated personal rule, not a property of every model
- also use external agents → if you have more than one agent CLI around, let it drive them non-interactively as workers. That’s also the natural place for a review from a different model family
- look at the results, not just delegate → this is what makes it orchestration instead of a routing table
That last one is really the whole thing. The value is that it looks at the subagents and their results and steers them, killing a bad approach after one round instead of five.
The word I keep coming back to, as weird or humanized as it sounds, is advisor. Orchestrator and advisor: it plans, it delegates, it judges, it advises. It can type, but typing is not why you put it in that seat.
Tell it explicitly
The important part is that you need to tell it.
Left alone, a capable model will helpfully do the work itself. That’s what it’s trained for: you gave it a task, it completes the task. It will not spontaneously decide that step four should go to something cheaper. It just gets on with it, competently, at the premium rate, including every search and every helper task on the way.
So don’t let it take over every helper step by accident. Tell it which role it has, which work can be delegated and which decisions it still owns.
Without that instruction the whole pattern just doesn’t happen.
When to skip the layer completely
Two-layer setups are not always right, and this is where I think people overbuild.
If you are doing a plan + implementation setup, the task is fuzzy and the approach is not settled. Somebody still has to decide what “done” even means. Then yes, take the strongest model for the plan/orchestration agent. That’s exactly the shape it is good at.
But if you already have well defined tasks, you as the human are basically the orchestrator. You did the decomposition, you know what has to happen, there is no judgment left to delegate. Then the orchestration layer is pure overhead, you are paying a model to forward instructions you already wrote. Call the worker model directly.
Keep the review either way. Skipping the orchestration layer means you took over the steering, it doesn’t mean the work stops needing a second pair of eyes from another family.
The task gets easier before the worker sees it
“Use cheaper models for helper tasks” was how I first explained the saving. That was only half the reason. The strong model also removes uncertainty. It decides what is actually wanted, what “done” means and which parts can become small, bounded worker tasks.
If a worker starts struggling, bring the problem back up. Re-plan it or escalate the model instead of burning more effort on a task that was never clear enough. That is where the accepted-result cost drops. I go deeper into that in A cheap run is not a cheap task.
Autonomous does not mean hands-off
One last boundary, because it is easy to look at 24 delegated tasks in one run and 30 agents and subagents in another and conclude that the number was the method.
It wasn’t. The run worked because the investigation had a measurement contract, and because I changed that contract when the evidence was too weak. The agents executed autonomously. I still decided what would count as proof.
You can read the complete example in The most valuable thing an agent did for me was not code.
That’s basically the whole idea: use the expensive model where its judgment changes the shape of the work. Let the cheaper models execute the shape it created. And keep your hands on the steering wheel.