OpenClaw Model Escalation Architecture: Routing Hard Work to a Bigger Model
An escalation path the agent does not know about is dead code with a nice comment on top.
In June I wrote a small shell script whose whole job was to send hard work from a Sonnet-class coding agent up to the strongest model I had access to. The header comment listed the categories it was for: pull request work, plus a few named product builds where a wrong call costs a week. The script worked. I tested it. Then for the rest of the summer the agent it was built for never called it once, and when the operator finally asked why that agent seemed weaker on hard problems than its peers, the answer was short and embarrassing.
The agent had never been told the script existed.
Two Places a Router Can Live
Most writing on model routing assumes the router sits in a gateway. A request arrives, a classifier or embedding lookup guesses how hard it is, and the gateway picks a model. Cascades are the popular variant: try the cheap model, check the output against a schema or a judge, and retry on the expensive one if the check fails. For single-shot API traffic (summarize this, extract that) this is a good design and I have nothing to add to it.
Agents break it. A coding agent working a contract makes forty tool calls across twenty minutes, and there is no schema for "did it reason well about the migration." By the time a judge could tell, you have paid for the whole cheap attempt and possibly merged it. The gateway also cannot see the thing that actually signals difficulty, which is the agent noticing halfway through that it is out of its depth.
So for long-running agents I put the decision inside the agent. It runs on the daily-driver model all day and calls out to a stronger one, deliberately, for a bounded subtask. The gateway still handles provider failover. That is availability, a separate concern, and mixing the two in one config is how you end up paying frontier prices because a cheap provider returned a 529.
The Dispatcher Is One File
Escalation does not need a framework. It needs a command the agent can call, with a fallback chain behind it. The version below (trimmed, with example categories) shells out to the Claude CLI in print mode so the escalated call gets a clean context window instead of inheriting the caller's forty turns of tool output:
#!/usr/bin/env bash
# escalate.sh usage: escalate.sh <category> <prompt-file>
# Categories: pr-review, schema-migration, prod-incident, client-deliverable
set -euo pipefail
category="$1"; prompt_file="$2"
chain=(claude-fable-5 claude-opus-4-8) # strongest first
log=~/.openclaw/logs/escalations.jsonl
for model in "${chain[@]}"; do
start=$(date +%s)
if out=$(claude --model "$model" -p "$(cat "$prompt_file")" 2>&1); then
printf '{"ts":%s,"caller":"%s","category":"%s","model":"%s","secs":%s}\n' \
"$start" "${AGENT_NAME:-unknown}" "$category" "$model" \
"$(( $(date +%s) - start ))" >> "$log"
printf '%s\n' "$out"
exit 0
fi
echo "escalate: $model failed, trying next" >&2
done
echo "escalate: every model in chain failed" >&2
exit 1The category argument is required on purpose. It forces the caller to say why it is spending the money, and it turns the log into something you can actually audit later.
The Prompt Is the Router
Here is where mine failed. The categories lived in the script's header comment, and the agent's operating instructions (a long briefing file it loads at boot) did not contain the script's name, the word escalate, or the name of the bigger model. An agent only reaches for tools it can see. It will happily reason its way through a hard architecture question on the smaller model, produce something plausible, and move on, because from its point of view no other option exists.
The fix is a rule in the instructions, written as concretely as a cron schedule:
## Escalation (read before any non-trivial task) You run on claude-sonnet-5. For these categories, draft the question, then run ~/bin/escalate.sh <category> <file> and use its answer: pr-review any PR touching auth, billing, or a migration schema-migration any change to a production table prod-incident a live site or gateway is down client-deliverable copy or code a named person will receive today Also escalate if you have retried the same failing approach twice. Do NOT escalate formatting, file moves, status reports, or lookups. Log the category honestly. Escalations are reviewed weekly.
Notice what it names. Specific kinds of work and one behavioral trigger (two failed retries), with an explicit list of what does not qualify. That last list matters about as much as the first, since an agent told only "escalate hard things" will either never do it or start escalating its own status updates.
Categories Beat Difficulty
Asking an agent to judge whether a task is hard is asking it to know what it does not know. Models are bad at this, and smaller models are worse, which is backwards from what you need: the agent most in need of escalating is least equipped to notice.
Categories sidestep that. "Does this PR touch a migration" is a question a Sonnet-class model answers correctly nearly every time. Pick categories by blast radius, meaning what a wrong answer costs you, and let the agent match against them. The retry trigger then catches the cases your categories missed.
Zero Is an Alarm
If the escalation log shows no entries for a whole week from an agent that did real work that week, the path is unwired, and you should go read that agent's instructions before you trust anything else it shipped.
Test the Chain Live
The same script carried a second stale fact. Its header said the top model had been pulled from the market, written during a brief outage and never revisited. Anyone reading it (human or agent) would assume the chain started at the fallback. A one-line live call proved the top model answered fine:
claude --model claude-fable-5 -p "reply with the single word ok"
Run that for every model in every chain on a schedule, and write the result somewhere agents read, ideally the fleet registry with a verified timestamp. Comments about model availability rot faster than any other kind of comment I keep. Check the launcher while you are there, too. The registry article covers a runner that hardcoded --model claude-sonnet-5 in two places while the docs claimed an upgrade, and an escalation rule written against the wrong base model reads very differently.
What It Costs
The model selection numbers on this site make the case for the whole pattern. When I moved my main agent from Opus to Sonnet for a week, daily cost fell about 70% (from $85 to $25) and the operator had to clarify questions three times as often. Neither side of that trade is acceptable as a permanent setting. Escalation is how you get most of the savings without most of the clarifying.
Each escalated call costs more than the same call on the base model, but it carries a fresh, small context instead of a session's accumulated history, so the bill is closer to a short Opus question than a long Opus session. Watch the ratio weekly. If one category climbs past a handful of calls a day, either the category is too broad or that agent belongs on the bigger model full time, and the cost architecture guide has the budget math for that decision.
Internal Links & Further Reading
- Model Selection Strategy: When to Use Opus, Sonnet, Flash, and DeepSeek →
The static, per-role assignment that escalation sits on top of.
- OpenClaw Fleet Registry Architecture →
Where verified model availability and launcher flags should live.
- OpenClaw Context Window Architecture →
Why the escalated call should start with a clean context.
- OpenClaw Error Handling and Recovery Patterns →
Provider failover and retries, kept separate from escalation.
FAQ
Q: Why not run every agent on the strongest model?
Because most of an agent's day is file moves and lookups, and paying frontier rates for those buys nothing. Escalation keeps the expensive model for the small share of work where a wrong answer is costly.
Q: Should the escalated model take over the rest of the session?
Usually no. Send it one bounded question with the context it needs, take the answer, and let the daily driver carry on. A permanent handoff makes sense only when the whole remaining task falls in an escalation category, and then it is cleaner to requeue the task to an agent that runs on the bigger model.
Q: How do I know escalation is helping?
Read the log next to outcomes. Pick ten escalated PR reviews and ten that were not escalated, and compare how many needed follow-up fixes. It is a small sample, and it will still tell you more than any dashboard about whether your categories are right.
The Bottom Line
Put the escalation decision inside the agent, give it a one-file dispatcher with a tested fallback chain, and write the categories into the instructions the agent loads at boot. Name kinds of work, never vague difficulty.
Then check the log. An empty one means the agent has never heard of the script.
Skip the trial and error
Get the OpenClaw Starter Kit — config templates, 5 ready-made skills, deployment checklist. Everything you need to go from zero to running in under an hour.
$14 $6.99
Get the Starter Kit →Also in the OpenClaw store
Get the free OpenClaw deployment checklist
Production-ready setup steps. Nothing you don't need.