← All Articles
ArchitectureOperations•8 min read•Oct 6, 2026

OpenClaw Config Change Architecture: Predictions, Deadlines, and Edits That Undo Themselves

A gateway that restarts cleanly on a new config has proven it can parse JSON. That is all it has proven.

The gateway on my Mac mini starts from one file. The LaunchAgent passes --config /Users/jkw/.openclaw/openclaw.json and everything else follows from what is in there: the port, the auth token, which tools each agent may call, which paths it may read. Change one line and you have changed the behaviour of every agent in the fleet. The failure I designed around this time has nothing to do with typos. You unload the gateway to edit that file, the self-healing watchdog notices sixty seconds later that the job is gone, and it loads the gateway back with your edit half finished.

Backups would not have helped.

What the Backup Guides Get Right

Search for how to change openclaw.json safely and the top guides agree on a routine. Stop the gateway. Copy the file somewhere with a date in the name. Make the change, start the gateway, run a health check, and if it fails, copy the old file back and restart again. The longer guides add a reminder that the file holds API keys and should never land in a public repo, which is correct and worth repeating.

I follow all of that. The trouble is the step where the health check passes, because a passing health check is where those guides end and where my real failures began. The gateway answers on port 8080. It also, quietly, does something you did not intend, and the backup sits in a folder getting older until you have forgotten which change it was protecting you from.

Write the Prediction First

Every config change is a claim about the future. "After this edit, the mira-alexandra agent can no longer run shell commands." "After this edit, agents default to the cheaper model." If you cannot say what will be observably different, you do not understand the change well enough to make it, and if you can say it, you can check it.

So before any edit to a system file, I write a small manifest next to it. It names the file, states the prediction in a sentence, gives a probe command that exits zero only if the prediction came true, and sets a deadline. Mine look like this:

# changes/2026-10-06-alexandra-no-exec.yaml
file: ~/.openclaw/openclaw.json
why: alexandra workspace should never shell out; read/write/edit only
prediction: >
  security.agents.mira-alexandra.allowedTools no longer contains exec,
  and a test request asking that agent to run a command is refused.
probe: ~/.openclaw/probes/alexandra-no-exec.sh
deadline_minutes: 10
on_fail: revert

The probe is the hard part to write, and the part that matters. A probe that reads the config file back and greps for the new value is nearly worthless, since it confirms that you saved the file. A good probe asks the running system. For a tool restriction, it sends the agent a harmless request that needs the tool and expects a refusal. For a model change, it reads the argv of the running process.

I learned that last one the expensive way, and the fleet registry article tells the whole story: a status doc said an agent had been upgraded on June 30, while the launcher hardcoded --model claude-sonnet-5 in two places and won every time. Any probe that read the doc would have passed for months.

An Apply Script With a Deadline

The manifest is only paperwork until something enforces it. This script takes a manifest and a candidate config, and it is the only way config reaches the live path on my machine:

#!/bin/bash
# apply-config.sh <manifest.yaml> <candidate.json>
set -u
manifest="$1"; candidate="$2"
live=~/.openclaw/openclaw.json
label=com.openclaw.gateway
domain="gui/$(id -u)"
hold=~/.openclaw/selfheal/$label/hold
probe=$(yq '.probe' "$manifest")
deadline=$(( $(date +%s) + 60 * $(yq '.deadline_minutes' "$manifest") ))

jq empty "$candidate" || { echo "REJECT not valid JSON"; exit 1; }

touch "$hold"                       # watchdog: hands off until we finish
stamp=$(date +%Y%m%dT%H%M%S)
cp "$live" "$live.pre-$stamp"
cp "$candidate" "$live.tmp" && mv "$live.tmp" "$live"   # atomic swap
launchctl kickstart -k "$domain/$label"

ok=0
while [ "$(date +%s)" -lt "$deadline" ]; do
  if curl -fsS -m 5 http://localhost:8080/health >/dev/null && bash $probe; then
    ok=1; break
  fi
  sleep 15
done

if [ "$ok" -eq 1 ]; then
  echo "$stamp KEPT $manifest" >> ~/.openclaw/changes.log
elif [ "$(yq '.on_fail' "$manifest")" = keep ]; then
  echo "$stamp UNPROVEN $manifest" >> ~/.openclaw/changes.log
else
  cp "$live.pre-$stamp" "$live"
  launchctl kickstart -k "$domain/$label"
  echo "$stamp REVERTED $manifest" >> ~/.openclaw/changes.log
fi
rm -f "$hold"
exit $(( 1 - ok ))

The hold file matters more than anything else in there. It is the same file the watchdog in the self-healing architecture already respects, and touching it before the swap closes the half-edited reload described at the top. The atomic mv closes it a second time. The gateway can never read a file that is partly old and partly new, because the live path points at one complete file or the other.

Then the default. Notice that the script reverts unless the probe passes. Most rollback setups do the opposite and revert only when something visibly breaks, which means a change that silently fails to take effect gets kept forever. Here, silence counts as failure. If I wrote a probe that cannot pass, I find out within ten minutes, with the old config back in place, instead of finding out in three weeks when an agent does the thing I thought I had forbidden.

The Log Is the Changelog

changes.log gets one line per attempt, and the manifests stay in a folder under version control (without the config itself, for the key reasons above). Six months on, that folder answers the question nobody can answer from a backup directory full of timestamped JSON: why does this line say what it says? The why field holds a sentence somebody wrote while they still remembered.

A REVERTED line is useful too. Three reverts in a row on the same manifest probably means the prediction is wrong rather than the edit, the probe is testing something the change was never going to affect, and I go and rewrite the probe before touching the config again.

When an Agent Writes the Change

Agents can write to these files as easily as I can, and an agent that hits a rate limit mid-task is very tempted to raise it. I would rather they did that through the apply script than by hand, and the manifest requirement turns out to be a decent test of whether the agent understood its own change. An agent that cannot write a falsifiable prediction for an edit has usually guessed at the edit.

One category stays with a person. Anything under security that widens access, a new tool in an allowedTools list or a broader pathRestrictions glob, goes through an approval gate before the apply script ever sees it. Narrowing access can go straight through, since the worst outcome is an agent that cannot do something and complains. The human-in-the-loop guide covers how that gate is built.

Generated files are the other exception, and they get refused outright. The fleet-state markdown and the skill-routing block inside each CLAUDE.md both carry a do-not-edit header, and the apply script rejects any manifest pointing at them. The right change there is to the source, then a re-render.

The Config You Cannot Probe

Some settings only matter at 3 a.m. under load, and for those the honest manifest says so, sets on_fail: keep, and names the day somebody will go and look.

Internal Links & Further Reading

FAQ

Q: Why not keep openclaw.json in git and roll back with git?

Because the file holds the gateway auth token and node tokens, and a repo full of secrets is a worse problem than a missing rollback. Version the manifests and the probes. Keep the config, and its pre-change copies, on the machine with tight permissions.

Q: Isn't ten minutes a long time for the gateway to run an unproven config?

For most edits the probe passes in the first fifteen-second loop and the wait never happens. The deadline exists for probes that need a cron tick or a real message to arrive. Set it per manifest, as short as the probe allows.

The Bottom Line

Back up before every change, like every guide says. Then write down what the change should do, give it a probe that asks the running gateway, and let a script revert anything that cannot prove itself before the deadline.

Pause the watchdog while you work. It is trying to help, and it will.

Get the free OpenClaw deployment checklist

Production-ready setup steps. Nothing you don't need.