← All Articles
ArchitectureReliability•9 min read•Oct 4, 2026

OpenClaw Self-Healing Architecture: Supervisors, Watchdogs, and the Evidence a Restart Destroys

A restart that works perfectly is also the fastest way to make sure you never learn why the process died.

My gateway and the Telegram bridge have been killed together more than once, both with SIGTERM, both within the same minute, and launchd brought both back before anyone was awake to look. That felt like a win for about a week. Then it happened again, and I had nothing to show the operator except two log lines saying the processes had exited with status 143 and a fresh PID on each. The supervisor had healed the system and erased the crime scene in the same motion.

Recovery is the easy half.

The Usual Three Tiers

Most self-healing guides for OpenClaw describe the same ladder, and the ladder is fine. Tier one is the service manager itself: a LaunchAgent with KeepAlive on macOS, or Restart=always under systemd, which respawns a dead process within seconds. Tier two is a watchdog on a timer that runs a health check and restarts the gateway when the check fails. Tier three is escalation, which in the fancier setups means handing the logs to a model for diagnosis and in every setup should mean a message to a human.

I run that ladder. This article is about three things the ladder usually leaves out: the failure states KeepAlive cannot see, the evidence you lose by restarting, and how you prove the healer itself is safe to leave running. The gateway architecture guide has the base plist, and I will assume yours looks roughly like it.

What KeepAlive Cannot See

KeepAlive watches one thing, which is whether the process launchd started is still alive. A gateway can be broken in at least three ways that have nothing to do with that:

  • Not registered. Something (an updater, a botched launchctl bootout, you at midnight) unloaded the job. launchd is not supervising a process it no longer knows about, so nothing will ever restart it.
  • Registered but not running. The job is loaded and launchd has stopped trying, usually because the process crashed fast enough, often enough, that launchd throttled it. launchctl print will show the job with no PID.
  • Running but not answering. The process is alive, holding its port, and stuck. From launchd's point of view this is a perfectly healthy service.

The first one is the nasty one, because the log tail you would normally reach for says nothing at all. A missing job does not write errors.

A Watchdog That Names the State

The watchdog runs as its own LaunchAgent with StartInterval set to 60, and it checks the three states in order. Each state gets a name, because a log line reading registered_not_running is something you can count across a month and a line reading "restarted gateway" is not.

#!/bin/bash
# gateway-selfheal.sh  (own LaunchAgent, StartInterval 60)
label=com.openclaw.gateway
domain="gui/$(id -u)"
plist=~/Library/LaunchAgents/$label.plist
state=~/.openclaw/selfheal/$label
mkdir -p "$state"
notify() { echo "$(date -u +%FT%TZ) $1" >> ~/.openclaw/alerts.log; }

# The operator stopped it on purpose. Leave it alone.
[ -f "$state/hold" ] && exit 0

if ! info=$(launchctl print "$domain/$label" 2>/dev/null); then
  reason=not_registered
elif ! grep -q 'state = running' <<<"$info"; then
  reason=registered_not_running
elif ! curl -fsS -m 5 http://localhost:8080/health >/dev/null; then
  reason=running_not_answering
else
  date +%s > "$state/last_ok"   # healthy: touch one file, say nothing
  exit 0
fi

# Circuit breaker: three recoveries in an hour means a human looks.
recent=$(find "$state" -name 'heal-*' -mmin -60 | wc -l)
if [ "$recent" -ge 3 ]; then
  notify "selfheal gave up on $label ($reason)"
  exit 1
fi

now=$(date +%s)
touch "$state/heal-$now"
launchctl print "$domain/$label" > "$state/before-$now.txt" 2>&1
launchctl bootout "$domain/$label" 2>/dev/null
launchctl bootstrap "$domain" "$plist"
launchctl kickstart -k "$domain/$label"
notify "selfheal recovered $label ($reason)"

Two lines in there carry most of the design. The hold file is how a deliberate stop stays stopped. Without it the watchdog and the operator fight, the operator unloads the gateway to edit openclaw.json, and sixty seconds later the watchdog loads it back with the half-edited config. The before- snapshot is the other one. It saves launchd's own view of the job, including the last exit status, before the bootout wipes it.

Use whole-second timestamps everywhere in a script like this. My first version compared a fractional timestamp inside bash integer arithmetic, and bash refused the comparison on the very first real run, which is a bad moment for a recovery script to discover a syntax problem.

Photograph the Room Before Anyone Cleans It

The watchdog catches what KeepAlive misses. Neither of them tells you why the gateway died, and for SIGTERM in particular the reason is almost never inside the gateway. Something outside sent it. On my Mac mini the pattern behind every cluster has been the same one described in the work queue article: the per-user process count climbed toward its ceiling, spawns started failing with EAGAIN, and long-running services got killed in the scramble.

I only know that because of a small wrapper. launchd starts the wrapper, the wrapper starts the real process, and when TERM arrives the wrapper writes down what the machine looked like at that instant before passing the signal along:

#!/bin/bash
# sigterm-wrapper.sh <label> <command> [args...]
label="$1"; shift
evdir=~/.openclaw/evidence/$label
mkdir -p "$evdir"

snapshot() {
  {
    echo "event=$1 child=$child"
    echo "uid_procs=$(ps -U "$(id -u)" -o pid= | wc -l)"
    echo "all_procs=$(ps -A -o pid= | wc -l)"
    memory_pressure | tail -1
    ps -U "$(id -u)" -o comm= | sort | uniq -c | sort -rn | head -15
  } > "$evdir/$(date +%Y%m%dT%H%M%S)-$1.txt" 2>&1
}

"$@" &
child=$!
trap 'snapshot TERM; kill -TERM "$child"; wait "$child"; exit 143' TERM
wait "$child"; rc=$?
[ "$rc" -ne 0 ] && snapshot "exit-$rc"
exit "$rc"

A shell trap cannot tell you which process sent the signal. It can tell you how crowded the room was, and the top-fifteen list of process names by count is usually enough to point at the culprit. Mine kept pointing at orphaned Node servers with no listening port, plus whatever sub-agent burst had just fired. Point ProgramArguments in the plist at the wrapper, with the gateway command as its trailing arguments. Nothing else changes.

That evidence is also what turned a vague worry into a number. Once the snapshots showed the user-account process count was the signal that tracked every kill (and total process count mostly measured which desktop apps were open), the fix stopped being "restart faster" and became an admission gate that refuses new sub-agent spawns above 550.

The Watchdog Needs a Pulse Too

Who restarts the watchdog is a fair worry, and the honest answer is that nothing does, which is why it writes last_ok on every healthy pass. Whatever already reads your heartbeats should read that file too and complain when it is more than a few minutes old. A dead watchdog and a healthy gateway look identical from the outside until the gateway breaks.

Test the Healer on Both Sides

Every self-healing writeup I read before building this tested recovery. Break the gateway, watch it come back, done. That proves half the contract. The other half is that the healer does nothing at all to a healthy service, and in my experience that half fails more often, because the health check has a typo, or the grep matches the wrong line in launchctl print, and now you have a script that restarts a working gateway every sixty seconds forever.

Run two tests before loading the watchdog for real:

  1. Quiet no-op. Point it at the real, healthy gateway and run it by hand five times. The only file that should change is last_ok. No heal- files, nothing in the alerts log, the gateway PID unchanged.
  2. Full recovery on a dummy. Make a throwaway LaunchAgent (label com.openclaw.selfheal-dummy, program /bin/sleep 86400), point a copy of the watchdog at it with the curl check removed, then bootout it and confirm the next pass brings it back. Never rehearse a bootout on the real gateway while agents are mid-task.

Then trip the circuit breaker on the dummy deliberately, so you have seen the "gave up" alert arrive once while nothing was on fire.

Where AI Diagnosis Fits

Some setups put a model in tier three and let it read logs and propose repairs. I have no objection, with one condition: give it the evidence directory and the before- snapshots, never just the gateway log. A model reading a log that ends in an exit code will invent a cause. A model reading a snapshot that shows hundreds of processes owned by the agent account has something real to reason about, and any repair it proposes should still go through the same approval gates as any other change to a production service.

Internal Links & Further Reading

FAQ

Q: Should the watchdog run more often than once a minute?

Probably not. KeepAlive already handles the fast case of a process dying. The watchdog exists for states that persist, and a minute of detection delay costs far less than a check that runs every few seconds and occasionally races launchd's own restart.

Q: Does the SIGTERM wrapper slow down shutdown?

By the time it takes to run ps twice and memory_pressure once, which is well under a second on a Mac mini. If the gateway needs a drain period on TERM (the zero-downtime deployment guide uses one), the wrapper forwards the signal and waits, so the drain still happens.

The Bottom Line

Keep KeepAlive, add a watchdog that names which of the three states it found, and wrap the process so a SIGTERM leaves a snapshot behind. Give the operator a hold file. Put a circuit breaker on recovery.

And test that the healer leaves a healthy gateway alone, because that is the test nobody runs.

Get the free OpenClaw deployment checklist

Production-ready setup steps. Nothing you don't need.