← All Articles
ArchitectureOperations•8 min read•Oct 7, 2026

OpenClaw Backup and Restore Architecture: Recovery Classes, Quiet Mode, and the First Hour Back

A snapshot is a picture of the past. Restore it onto a running fleet and the fleet will try to live in it.

Here is the morning I design my backups around. The nightly snapshot of the Mac mini runs at 3 a.m. At 7 a.m. the digest.daily job runs, writes a summary, and sends it to Telegram. At 9 the SSD dies. I put a replacement machine on the desk, restore the 3 a.m. copy, log in, and every LaunchAgent in ~/Library/LaunchAgents loads at once. In the restored queue, the digest job is still sitting there with status queued, because at 3 a.m. it was.

So it runs again.

What the Backup Guides Agree On

The top results for OpenClaw backup and restore are long, between two and three and a half thousand words, and they agree with each other on almost everything. Archive ~/.openclaw. Encrypt the archive, since openclaw.json holds API keys and channel tokens. Keep one copy somewhere other than the machine it describes. Verify the archive after writing it, and run the backup on a schedule with pruning so the folder does not eat the disk. The better ones add a restore drill on a clean machine, and one of them says a backup you have never restored is a hope, which I agree with completely.

What none of them cover is what the restored system does in its first hour. They stop at "the gateway is up and the health check passes." That is the moment an agent fleet is most dangerous, because the queue, and every watchdog sitting on top of it, wakes up believing it is still 3 a.m.

Sort Files by How They Come Back

"Back up everything" treats a memory file and a generated status page as the same kind of thing. They are opposites. I sort what lives on the machine by how it should return after a disaster, and each class gets different handling.

State nobody can recreate. MEMORY.md, the daily logs, and the SQLite files under ~/.openclaw/data. This is the only class that justifies most of the effort. The Mac mini setup guide describes the arrangement: Time Machine for the whole disk, plus a nightly cron job that commits the day's memory and logs to a private repository. One correction to the cron in the database-backed agents article. A plain cp crm.db taken while an agent is mid-write can copy a file that SQLite will later call corrupt, and in WAL mode it can miss committed rows still sitting in the -wal file. Use sqlite3 crm.db ".backup 'crm-copy.db'" or the backup API from the same article. The cp line works most nights, which is exactly why nobody notices it until the night it does not.

Secrets. The gateway auth token and node tokens live in openclaw.json, so that file goes into the encrypted archive and nowhere else. An older article on this site suggests keeping all the JSON configs in a private GitHub repo. I would not, for the reasons in the config change piece. Version the change manifests and probes in git. Keep the file with the tokens on the machine and in the encrypted copy.

Generated files. Do not restore these at all. My fleet-state markdown is rendered from the registry, and the skill-routing block in each project's CLAUDE.md is rendered from a YAML file. Restoring last night's render gives you a document that agrees with last night's registry. Restore the source, then re-render, and the output agrees with whatever the source says now. The fleet registry article explains why that render step exists in the first place.

The work queue. This one gets restored, and then distrusted. It is the subject of the rest of this page.

Why the Queue Lies After a Restore

The work queue design protects against duplicates with a dedupe_key and a unique index, so two producers asking for digest.daily:2026-09-30 collapse into one job. That protection lives inside the database. The restored database has never heard of the 7 a.m. send, because the send happened after the snapshot, and the message itself now sits in a Telegram chat the queue cannot see.

Leases make it worse in a quieter way. Any job that was running at 3 a.m. comes back with a lease_until in the past and its attempts counter one higher. To the dispatcher this looks exactly like a worker that crashed, and reclaiming expired leases is the whole point of leases. The job runs from the top. If it was a content contract, you get a second branch. If it was an outbound email, somebody gets it twice.

The supervisors do their part too. A self-healing watchdog that finds a registered job not running will start it, which is what it is for, and on a restored machine it fires for every label at roughly the same moment. That is the same burst of spawns the work queue article opens with, arriving on a machine you have owned for twenty minutes.

Restore Into Quiet Mode

My restore script never brings the fleet up live. It lays the files down, puts every supervised job on hold, and sets a flag the dispatcher reads before admitting any job whose kind has an outbound side effect. Jobs leased at snapshot time get pulled aside for a person to look at.

#!/bin/bash
# restore.sh <archive.tar.age>
set -eu
age -d -i ~/.config/age/restore.key "$1" | tar -x -C ~/

# 1. nothing supervised starts on its own
for plist in ~/Library/LaunchAgents/com.openclaw.*.plist; do
  label=$(basename "$plist" .plist)
  mkdir -p ~/.openclaw/selfheal/$label
  touch ~/.openclaw/selfheal/$label/hold
done

# 2. dispatcher admits local work only
touch ~/.openclaw/QUIET

# 3. anything mid-flight at snapshot time waits for review
sqlite3 ~/.openclaw/data/queue.db "
  UPDATE jobs SET status = 'needs_review'
  WHERE status = 'running'
     OR (status = 'queued' AND run_after < unixepoch());"

# 4. regenerate, never restore, rendered files
~/.openclaw/bin/render-fleet-state
~/.openclaw/bin/render-skill-routing

echo "restored in QUIET mode; review needs_review jobs, then rm ~/.openclaw/QUIET"

Step 3 is broad on purpose. A queued job whose run_after passed between the snapshot and the restore might have run on the dead machine, or might not have. The queue cannot tell, and I would rather read a short list of job names over coffee than guess. Each one gets checked against the outside world (the Telegram chat for anything that messages me, the git remote for anything that writes code) and then either requeued or closed.

Inside the dispatcher, quiet mode is three lines in the admission check: if ~/.openclaw/QUIET exists and the job kind declares an outbound resource, refuse with reason: "quiet" and a long retry. Local work like embedding backfills still runs, which is handy, because it warms the machine while you sort out the rest.

The Telegram Bridge Comes Back Last

It is the one part of the fleet that talks to a human who has no idea a restore happened.

Drill It on the Wrong Machine

The guides that recommend restore drills say to use a clean environment. Go further and use a machine that differs in some way you did not choose. A different username is the best one. Hardcoded paths like /Users/jkw/.openclaw/openclaw.json in a LaunchAgent plist are invisible on the original box and break immediately on a box with any other home directory. A drill on an identical clone finds none of that.

Time the drill, and stop the clock when the fleet is doing real work again, after the review list is cleared and quiet mode is off. Gateway-answering-on-8080 is a much earlier and much less honest number.

Then run the same probes from the config change article against the restored machine. Read the model from the argv of the running process, ask a restricted agent to do something it should refuse, and compare both against the registry. A restore that passes its health check and quietly runs the wrong model is the outcome I worry about most, since the backup faithfully preserved whatever was wrong before.

Internal Links & Further Reading

FAQ

Q: Does Time Machine alone count as an OpenClaw backup?

For the disk, yes. For the SQLite files, only if nothing was writing when it ran, so pair it with a real .backup of each database. And remember that a Time Machine restore brings back the queue too, with all the problems above.

Q: Why not skip restoring the queue and start empty?

You lose every job that was waiting, and some of those were real requests from a person. Restoring it and marking the uncertain rows for review keeps the requests while refusing to act on them blind.

The Bottom Line

Back up memory and SQLite with tools that understand SQLite, and keep the token-bearing config out of git. Then restore into quiet mode, with every watchdog on hold and every half-finished job waiting for a human to check what already happened.

The snapshot does not know what time it is. You do.

Get the free OpenClaw deployment checklist

Production-ready setup steps. Nothing you don't need.