← All Articles
ArchitectureCost•9 min read•Oct 3, 2026

OpenClaw Model Escalation Architecture: Routing Hard Work to a Bigger Model

An escalation path the agent does not know about is dead code with a nice comment on top.

In June I wrote a small shell script whose whole job was to send hard work from a Sonnet-class coding agent up to the strongest model I had access to. The header comment listed the categories it was for: pull request work, plus a few named product builds where a wrong call costs a week. The script worked. I tested it. Then for the rest of the summer the agent it was built for never called it once, and when the operator finally asked why that agent seemed weaker on hard problems than its peers, the answer was short and embarrassing.

The agent had never been told the script existed.

Two Places a Router Can Live

Most writing on model routing assumes the router sits in a gateway. A request arrives, a classifier or embedding lookup guesses how hard it is, and the gateway picks a model. Cascades are the popular variant: try the cheap model, check the output against a schema or a judge, and retry on the expensive one if the check fails. For single-shot API traffic (summarize this, extract that) this is a good design and I have nothing to add to it.

Agents break it. A coding agent working a contract makes forty tool calls across twenty minutes, and there is no schema for "did it reason well about the migration." By the time a judge could tell, you have paid for the whole cheap attempt and possibly merged it. The gateway also cannot see the thing that actually signals difficulty, which is the agent noticing halfway through that it is out of its depth.

So for long-running agents I put the decision inside the agent. It runs on the daily-driver model all day and calls out to a stronger one, deliberately, for a bounded subtask. The gateway still handles provider failover. That is availability, a separate concern, and mixing the two in one config is how you end up paying frontier prices because a cheap provider returned a 529.

The Dispatcher Is One File

Escalation does not need a framework. It needs a command the agent can call, with a fallback chain behind it. The version below (trimmed, with example categories) shells out to the Claude CLI in print mode so the escalated call gets a clean context window instead of inheriting the caller's forty turns of tool output:

#!/usr/bin/env bash
# escalate.sh  usage: escalate.sh <category> <prompt-file>
# Categories: pr-review, schema-migration, prod-incident, client-deliverable
set -euo pipefail
category="$1"; prompt_file="$2"
chain=(claude-fable-5 claude-opus-4-8)   # strongest first
log=~/.openclaw/logs/escalations.jsonl

for model in "${chain[@]}"; do
  start=$(date +%s)
  if out=$(claude --model "$model" -p "$(cat "$prompt_file")" 2>&1); then
    printf '{"ts":%s,"caller":"%s","category":"%s","model":"%s","secs":%s}\n' \
      "$start" "${AGENT_NAME:-unknown}" "$category" "$model" \
      "$(( $(date +%s) - start ))" >> "$log"
    printf '%s\n' "$out"
    exit 0
  fi
  echo "escalate: $model failed, trying next" >&2
done
echo "escalate: every model in chain failed" >&2
exit 1

The category argument is required on purpose. It forces the caller to say why it is spending the money, and it turns the log into something you can actually audit later.

The Prompt Is the Router

Here is where mine failed. The categories lived in the script's header comment, and the agent's operating instructions (a long briefing file it loads at boot) did not contain the script's name, the word escalate, or the name of the bigger model. An agent only reaches for tools it can see. It will happily reason its way through a hard architecture question on the smaller model, produce something plausible, and move on, because from its point of view no other option exists.

The fix is a rule in the instructions, written as concretely as a cron schedule:

## Escalation (read before any non-trivial task)
You run on claude-sonnet-5. For these categories, draft the question,
then run ~/bin/escalate.sh <category> <file> and use its answer:
  pr-review          any PR touching auth, billing, or a migration
  schema-migration   any change to a production table
  prod-incident      a live site or gateway is down
  client-deliverable copy or code a named person will receive today
Also escalate if you have retried the same failing approach twice.
Do NOT escalate formatting, file moves, status reports, or lookups.
Log the category honestly. Escalations are reviewed weekly.

Notice what it names. Specific kinds of work and one behavioral trigger (two failed retries), with an explicit list of what does not qualify. That last list matters about as much as the first, since an agent told only "escalate hard things" will either never do it or start escalating its own status updates.

Categories Beat Difficulty

Asking an agent to judge whether a task is hard is asking it to know what it does not know. Models are bad at this, and smaller models are worse, which is backwards from what you need: the agent most in need of escalating is least equipped to notice.

Categories sidestep that. "Does this PR touch a migration" is a question a Sonnet-class model answers correctly nearly every time. Pick categories by blast radius, meaning what a wrong answer costs you, and let the agent match against them. The retry trigger then catches the cases your categories missed.

Zero Is an Alarm

If the escalation log shows no entries for a whole week from an agent that did real work that week, the path is unwired, and you should go read that agent's instructions before you trust anything else it shipped.

Test the Chain Live

The same script carried a second stale fact. Its header said the top model had been pulled from the market, written during a brief outage and never revisited. Anyone reading it (human or agent) would assume the chain started at the fallback. A one-line live call proved the top model answered fine:

claude --model claude-fable-5 -p "reply with the single word ok"

Run that for every model in every chain on a schedule, and write the result somewhere agents read, ideally the fleet registry with a verified timestamp. Comments about model availability rot faster than any other kind of comment I keep. Check the launcher while you are there, too. The registry article covers a runner that hardcoded --model claude-sonnet-5 in two places while the docs claimed an upgrade, and an escalation rule written against the wrong base model reads very differently.

What It Costs

The model selection numbers on this site make the case for the whole pattern. When I moved my main agent from Opus to Sonnet for a week, daily cost fell about 70% (from $85 to $25) and the operator had to clarify questions three times as often. Neither side of that trade is acceptable as a permanent setting. Escalation is how you get most of the savings without most of the clarifying.

Each escalated call costs more than the same call on the base model, but it carries a fresh, small context instead of a session's accumulated history, so the bill is closer to a short Opus question than a long Opus session. Watch the ratio weekly. If one category climbs past a handful of calls a day, either the category is too broad or that agent belongs on the bigger model full time, and the cost architecture guide has the budget math for that decision.

Internal Links & Further Reading

FAQ

Q: Why not run every agent on the strongest model?

Because most of an agent's day is file moves and lookups, and paying frontier rates for those buys nothing. Escalation keeps the expensive model for the small share of work where a wrong answer is costly.

Q: Should the escalated model take over the rest of the session?

Usually no. Send it one bounded question with the context it needs, take the answer, and let the daily driver carry on. A permanent handoff makes sense only when the whole remaining task falls in an escalation category, and then it is cleaner to requeue the task to an agent that runs on the bigger model.

Q: How do I know escalation is helping?

Read the log next to outcomes. Pick ten escalated PR reviews and ten that were not escalated, and compare how many needed follow-up fixes. It is a small sample, and it will still tell you more than any dashboard about whether your categories are right.

The Bottom Line

Put the escalation decision inside the agent, give it a one-file dispatcher with a tested fallback chain, and write the categories into the instructions the agent loads at boot. Name kinds of work, never vague difficulty.

Then check the log. An empty one means the agent has never heard of the script.

Get the free OpenClaw deployment checklist

Production-ready setup steps. Nothing you don't need.