← All Articles
ArchitectureScaling•12 min read•Sep 30, 2026

OpenClaw Work Queue Architecture: Leases, Concurrency Limits, and Backpressure

Every agent fleet grows a queue eventually. The only choice is whether you design it or discover it at 3 a.m. in a process table.

The first time my fleet fell over, nothing was wrong with any single agent. A morning digest, a failure-replay job, and a content contract all fired within the same two minutes, each one spawned its own sub-agents, and each sub-agent spawned a shell, a Node process, and a browser. The Mac mini hit its per-user process ceiling. New spawns failed with EAGAIN, the gateway got a SIGTERM it did not deserve, and the Telegram bridge went down with it.

Nobody had asked for that much work. Everyone had simply started it.

Cron Should Only Be a Trigger

Most OpenClaw setups start with cron jobs that launch agents directly. That works until two things happen at once. Cron has no idea what else is running, how much memory is free, or whether the job it launched yesterday is still going, so it will cheerfully start a second copy of a job whose first copy is stuck on a hung API call.

The fix is a layer between "something wants work done" and "an agent process starts." Cron, webhooks, Telegram messages, and other agents all become producers. They write a job row and walk away. A small dispatcher owns the decision about when that job actually runs, on which worker, and whether it runs at all today.

producers                       queue (sqlite / postgres)            workers
---------                       --------------------------           -------
cron  ─┐                        jobs: id, kind, payload,             agent runner A
webhook├──> enqueue(job) ────>        priority, dedupe_key,  <─lease─ agent runner B
agent ─┘                              lease_until, attempts          browser pool (2)
                                         │
                                  dispatcher: admission control,
                                  per-resource limits, reaper

You do not need Kafka for this. A single SQLite file with WAL mode handles a personal or small-team fleet comfortably, and it is the same database the database-backed agents guide already has you running. Postgres with SELECT ... FOR UPDATE SKIP LOCKED is the upgrade path once workers live on more than one machine.

Claim Work With Leases

A worker claims a job by setting lease_until to a time in the near future. While it works, it renews the lease on a heartbeat. If the worker dies, gets killed by the OS, or hangs forever waiting on a model provider, the lease simply expires and the job becomes claimable again. Nobody has to notice the crash for recovery to happen.

-- claim one job atomically (SQLite 3.35+)
UPDATE jobs
SET    lease_owner = :worker_id,
       lease_until = unixepoch() + 120,
       attempts    = attempts + 1
WHERE  id = (
  SELECT id FROM jobs
  WHERE  status = 'queued'
    AND  (lease_until IS NULL OR lease_until < unixepoch())
    AND  run_after <= unixepoch()
  ORDER BY priority DESC, created_at
  LIMIT 1
)
RETURNING id, kind, payload, attempts;

-- heartbeat every 30s while the agent runs
UPDATE jobs SET lease_until = unixepoch() + 120
WHERE  id = :id AND lease_owner = :worker_id;

The lease_owner check on the heartbeat matters more than it looks. Picture a worker that stalls for three minutes on a garbage-collection pause, loses its lease, and wakes up still believing it owns the job while a second worker has already picked it up. Without the owner check, both keep renewing and both write results. With it, the stale worker's heartbeat updates zero rows, which is its signal to stop and throw its work away.

Lease length is a tradeoff you set per job kind. Short leases recover faster. Long ones survive slow model calls without heartbeat traffic. I use two minutes with a thirty-second heartbeat for almost everything, and a separate kind with a ten-minute lease for the handful of jobs that hold a browser session open.

Limit Each Resource Separately

A global "max four concurrent jobs" setting is the obvious first move, and it is wrong in an instructive way. A job that only calls a model API costs almost nothing locally. A job that drives headless Chrome costs a few hundred megabytes and a dozen processes. Four of the first kind is idle. Four of the second on an 8 GB mini is swap.

Declare what each job kind consumes, then give each resource its own semaphore:

const RESOURCES = {
  browser:     { limit: 2 },   // headless Chrome sessions
  local_model: { limit: 1 },   // one Ollama generation at a time
  anthropic:   { limit: 6 },   // stay under org rate limits
  shell:       { limit: 4 },   // coding agents with a live worktree
};

const JOB_KINDS = {
  "site.qa":          { needs: ["browser", "anthropic"], lease_s: 600 },
  "digest.daily":     { needs: ["anthropic"],            lease_s: 120 },
  "embed.backfill":   { needs: ["local_model"],          lease_s: 120 },
  "content.contract": { needs: ["shell", "anthropic"],   lease_s: 900 },
};

The dispatcher only hands a job to a worker when every resource it needs has a free slot. Acquire them in a fixed order (alphabetical is fine) so two jobs never deadlock each holding half of what the other wants. Provider rate limits belong here too, since a 429 storm from six agents retrying at once is just backpressure arriving late and angry.

Admission Control: Saying No Early

Semaphores protect you from your own jobs. They do nothing about the rest of the machine. The owner has Chrome open with sixty tabs, a video call is running, and the process count is already near the ceiling before your fleet has spawned anything.

So the dispatcher checks the host before every claim. Pick one or two signals that actually correlate with your past failures and gate on those. For my fleet that turned out to be the process count for the user account the agents run under, not total processes and not load average. Total count mostly measured which desktop apps were open, and it kept blocking work on days when there was plenty of headroom.

async function admit(job) {
  const procs = await countProcesses({ uid: AGENT_UID });
  if (procs > 550) return { admit: false, reason: "uid_procs", retry_after_s: 60 };

  const freeMb = await freeMemoryMb();
  if (freeMb < 1024 && JOB_KINDS[job.kind].needs.includes("browser"))
    return { admit: false, reason: "low_memory", retry_after_s: 120 };

  return { admit: true };
}

Put the reason on the job row whenever admission refuses it. When someone asks why the daily digest ran forty minutes late, "deferred 38 times, uid_procs" is an answer. "It was slow" is not.

Backpressure Has to Reach the Producer

A queue that accepts everything and runs it eventually has only moved the problem. If a webhook fires two hundred times during an incident, you now have two hundred jobs, and the backlog will still be draining long after anyone cares about the results.

Three controls do most of the work, and they are not equally important. The first one matters most by a wide margin.

  • Dedupe keys. Every producer sets a dedupe_key (for example digest.daily:2026-09-30), and a unique index turns a duplicate enqueue into a no-op. Cron double-fires, webhook retries, and an agent that asks for the same report twice in one session all collapse into one job.
  • Per-kind queue caps. When site.qa already has twenty queued, the twenty-first enqueue fails loudly and the producer decides what to do.
  • Expiry. A job can carry stale_after. The morning briefing that could not run by noon should be dropped with a note. Delivering it at 4 p.m. helps nobody.

When an agent is the producer, return the refusal as a structured tool result it can reason about (the tool layer guide calls this the transient bucket, with a retry_after_s). Agents handle "queue full, try in ten minutes" surprisingly well. They handle a silent hang terribly.

Poison Jobs

Some jobs will never succeed. The payload references a file that was deleted, or the prompt reliably pushes the model into a loop that blows the lease every time. Leases make these worse on their own, because an expired lease looks exactly like a crashed worker and the job gets retried forever.

Cap attempts per kind (three is my default), back off exponentially between them with run_after, and move anything past the cap to a dead status with the last error attached. Then look at the dead list. A daily count of dead jobs by kind is one of the most useful numbers in the whole observability setup, because it catches the slow failures that never page anyone.

Sub-Agents Enqueue Too

If a parent agent can spawn children directly, your limits are decorative.

Route every spawn through the queue, including the ones an agent makes mid-task during multi-step workflows. Give child jobs a parent_id and a slightly higher priority than new top-level work, so a half-finished task completes before a fresh one starts. One more rule saves a lot of grief: a parent waiting on its children should release its own heavy resources while it waits, otherwise four parents can hold all four shell slots while their children sit in the queue behind them, and nothing moves until the leases expire.

Failure Modes Worth Knowing

The zombie that still writes

A worker loses its lease during a long pause, a second worker finishes the job, and then the first one wakes up and posts its own result to Telegram. The owner gets two slightly different digests.

Fix: check lease ownership on heartbeat and again right before every side effect. Pair it with the idempotency keys from the tool layer so a stale write is refused at the edge.

Orphans after a restart

The dispatcher restarts but the browser processes its workers launched keep running, holding memory and counting against the process gate. The gate then refuses new work because of processes nobody owns.

Fix: launch each worker in its own process group, record the group ID on the job row, and have the reaper kill groups whose lease has expired.

Priority starvation

Interactive requests from the owner always outrank batch work, which is right, until a busy week means the weekly backup job has not run in nine days.

Fix: age priorities. Add a point for every hour a job has waited so old batch work eventually outranks new chatter.

Start Smaller Than This

You can build everything above in an afternoon, and you probably should not build all of it on day one. The minimum useful version is a jobs table with a dedupe key, a lease column, and an attempts cap, plus one global limit on browser sessions. That alone would have prevented my process-limit outage.

Add per-resource semaphores when you have more than one expensive resource. Add host admission control the first time something else on the machine starves your agents, and use the signal that failure actually showed you rather than the one that sounds right. The performance and scaling guide covers what to measure before you tune any of these numbers.

Internal Links & Further Reading

To go deeper on the layers this article touches:

FAQ

Q: Do I need Redis or a message broker for an OpenClaw work queue?

Not for a single host or a small fleet. A SQLite table in WAL mode with an atomic claim query handles thousands of jobs a day, and you can inspect it with any SQL client. Move to Postgres with SKIP LOCKED when workers run on several machines. A dedicated broker earns its keep mostly at volumes most agent fleets never reach.

Q: How long should a job lease be?

Long enough to cover a slow model call plus one missed heartbeat. Two minutes with a thirty-second heartbeat suits most model-only jobs. Browser and coding jobs that hold external state for a while do better with a longer lease on their own job kind.

Q: What host signal should gate new agent spawns?

Whichever one tracked your last real outage. On a Mac mini that is often the per-user process count, since macOS enforces a per-user process limit and agent spawns fail with EAGAIN when it is reached. Free memory is the next candidate if you run browsers or local models. Load average tends to be too noisy to gate on.

Q: Should interactive messages go through the queue too?

Yes, at the highest priority, with admission control relaxed but not removed. An owner message that fails fast with "the machine is saturated, I will pick this up in a minute" is far better than a spawn that crashes the gateway for everyone.

The Bottom Line

Let producers ask for work and let one dispatcher decide when it runs. Claim with leases that expire on their own, limit each expensive resource separately, check the host before you admit anything, and cap retries so broken jobs die where you can see them. Route sub-agent spawns through the same door as everything else.

If your fleet has ever crashed from too much work that nobody asked for all at once, you already know which piece to build first.

Get the free OpenClaw deployment checklist

Production-ready setup steps. Nothing you don't need.