← All Articles
ArchitectureOperations•10 min read•Oct 1, 2026

OpenClaw Fleet Registry Architecture: One Source of Truth for Agents, Models, and Hosts

The status file your agents read every morning is either generated from evidence or slowly becoming fiction. There is no third state.

For most of August, the file my agents loaded at boot to learn who else was in the fleet was wrong. It listed an agent I had retired as active. It said the chief-of-staff agent ran on a model it had not used in weeks. Every agent that read it made small decisions on top of those errors (who to hand work to, which model to assume a peer could handle), and none of them had any reason to doubt it, because the file looked authoritative and was formatted beautifully.

Nobody lied. People edited it by hand.

Count the Copies

Pick any one fact about your fleet, say which model a given agent runs on, and go find every place it is written down. In my setup that fact lived in four places: the agent's profile config, the launcher script that actually starts the process, the shared fleet-state markdown file, and my own memory. They disagreed.

The worst case was the launcher. The status doc said one agent had been upgraded to a bigger model on June 30. The Python runner that launched it hardcoded --model claude-sonnet-5 in two separate places, and those two lines won every single time, because the process does what its argv says and nothing else. The doc was a statement of intent. Somebody wrote it the day of the decision and never checked that the change landed.

That is the normal failure, and it is quiet. A config change gets decided in a chat, written into a doc, half-applied to one file, and the gap survives until an agent behaves strangely enough that a human goes digging.

The Runtime Is the Authority

My rule now is blunt. A fact about the fleet is true only if you can read it off the thing that executes. The model is whatever the running process was launched with. The agent roster is whatever launchd (or systemd, or your supervisor) actually has loaded. Every doc and dashboard is derived from those, never the other way round.

This sounds obvious until you notice how many fleets run on the reverse. A wiki page decides what is supposed to be true, and the configs are expected to follow. They follow for about a week.

A Registry With Evidence Columns

The fix is a small database that stores facts along with how you know them. I use SQLite, the same file the database-backed agents setup already keeps around. Each row is one claim about one entity, and the evidence column is mandatory.

CREATE TABLE facts (
  entity      TEXT NOT NULL,   -- 'agent:cedar', 'host:mini-01', 'cron:daily-digest'
  attribute   TEXT NOT NULL,   -- 'model', 'persona', 'status', 'owner'
  value       TEXT NOT NULL,
  evidence    TEXT NOT NULL,   -- where this was read, in words a human can re-check
  probe       TEXT,            -- command that re-derives the value
  status      TEXT NOT NULL DEFAULT 'claimed',  -- 'claimed' | 'verified'
  verified_at INTEGER,
  PRIMARY KEY (entity, attribute)
);

A row starts as claimed. Anyone can claim things, including a human typing in a chat or an agent that just changed a config. It becomes verified only when its probe runs and returns the same value. A good evidence string reads like ~/.hermes/profiles/cedar/config.yaml model.default (verified 2026-08-23). A bad one says "upgraded, see chat." One of those you can check in ten seconds.

Probes should hit the most executable source available. Reading the YAML is fine when the launcher honors the YAML. When it does not, read the process itself:

# what the config claims
yq '.model.default' ~/.hermes/profiles/nova2/config.yaml

# what the running process was actually given
ps -o command= -p "$(pgrep -f nova-pty-runner)" | grep -o -- '--model [^ ]*'

When those two disagree you have found exactly the bug that bit me, and the registry should store the second value with the first one noted in evidence as the conflicting claim.

Render, Never Edit

The fleet-state file still exists. Agents still read it at boot, since a markdown table costs far fewer tokens than handing every agent a SQL client (the context window guide covers why that budget matters). What changed is who writes it. A render script queries verified rows only and overwrites the file, and the first lines of the output say so:

# Fleet State (generated from fleet registry)
<!-- Regenerated 2026-10-01 14:54 UTC by render_fleet_state.py.
     Source of truth: registry.db (verified rows only).
     Do NOT hand-edit. Change the registry, then re-render. -->

That comment is addressed to agents as much as to people. A coding agent asked to "update the fleet doc" will happily edit markdown unless the file tells it where the real lever is. Mine read the header and filed a registry claim instead. Run the renderer on a schedule and after every registry write, and put a pre-commit check on the file that rejects any diff not produced by the script.

Verification Decays

A fact verified in June tells you very little in October. Set a freshness window per attribute and let a nightly job re-run every probe. Model assignments I re-verify weekly, since a fleet-wide rollout (we moved six agents from one Gemini Flash version to the next on September 3) touches many rows at once. Host and persona facts change rarely and can go thirty days.

Anything past its window drops back to claimed and falls out of the rendered file. That feels aggressive the first time a row disappears. It is meant to. A missing row prompts a question, and a confidently wrong one prompts nothing.

One Writer Per Fact

I once had a JSON report that re-raised the same five resolved alerts every morning for more than seventeen days in a row, because the job that generated it rebuilt the file from scratch and erased the signal_resolved flags a different job had written the night before.

Two writers. No merge. Every day a third process (an agent, naturally) noticed, patched the flags back in, and wrote a note that the root cause was still open. The registry design removes this class of bug by giving each attribute exactly one owner. The prober owns status and verified_at. Whoever made a change owns the claim. Generated files have one writer, the renderer, and it never reads its own output back as input.

If you cannot avoid two jobs touching the same JSON, make both of them read, merge their own keys, and write. Never rebuild a shared file from scratch.

What Stays Out

Secrets stay out, obviously. Store the fact that an agent has a Telegram token and where it lives, never the token; the security architecture guide covers the vault side.

Fast-moving numbers stay out too. Process counts and queue depth change every few seconds and belong in metrics. Thresholds are different. The rule that says spawns halt above 550 processes for the agent user is a decision, it has an owner and a date, and it belongs in the registry where agents will find it.

Start With a Script and a Table

You do not need the full schema to get most of the value. Write one script that greps your agent configs and launcher files for model names, prints a markdown table, and stamps the header. Run it nightly. Delete the hand-written version the same day so nobody can keep editing it out of habit.

The first run will probably disagree with what you believe about your own fleet. Mine disagreed on two agents out of ten. That gap is the whole argument for building this, and it is a lot cheaper to find in a diff than in an agent that has been quietly routing hard problems to a model you thought you had replaced.

Internal Links & Further Reading

FAQ

Q: Can the registry just be a YAML file in git?

For a handful of agents, yes, as long as a script generates it from probes rather than a person typing it. Git history then doubles as your audit trail. Move to SQLite when you want freshness windows and per-attribute owners, which are awkward to express in a flat file.

Q: Should agents be allowed to write to the registry?

They should be allowed to write claims, and they should never be allowed to mark anything verified. Only the prober does that. An agent that just edited a config files a claim with the file path as evidence, and the next probe run confirms or rejects it.

Q: What if the registry and a rendered doc disagree?

The registry wins by definition, and the doc is regenerated. If the registry itself is wrong, fix the row with new evidence and let the doc correct itself on the next render.

The Bottom Line

Treat the running process as the authority on what your fleet is. Store every fact with the evidence that backs it, let a probe promote claims to verified, and generate every doc agents read from verified rows only. Give each fact one writer.

Then delete the hand-edited status file. You will miss it for a day.

Get the free OpenClaw deployment checklist

Production-ready setup steps. Nothing you don't need.