OpenClaw Multi-Tenant Architecture: Isolating Agents Per Client
How to run OpenClaw agents for several clients on one stack without leaking memory, credentials, or context between them.
The first cross-client leak I saw was a pricing number. An agent drafting a proposal for a small accounting firm quoted a discount tier that belonged to a completely different customer, a logistics company whose notes happened to live in the same memory directory, retrieved because the two documents both mentioned "annual renewal." Nobody got sued. The draft was caught in review. But it had passed every test we had, because none of our tests ever asked the only question that mattered: whose data is this?
That question gets expensive once you run agents for more than one client. Agencies hit it first. Anyone who has built a good internal agent and then thought "we could sell this" hits it eventually, usually about a week after the second customer signs.
The Leak Nobody Tests For
Single-tenant OpenClaw setups are built on a convenient assumption. Everything the agent can see belongs to one owner, so the agent can see everything. Memory files sit in one workspace. The retrieval index is one collection. Credentials live in one environment. The cron table is one list. Adding a second client to that design by creating a subfolder called clients/acme/ feels like isolation and behaves like a suggestion.
Cross-tenant contamination usually looks like a helpful agent doing its job well. Retrieval pulls the most relevant chunk, and relevance has no concept of ownership, so the best match for "renewal terms" might be another client's contract. A nightly summarizer compacts all of yesterday's notes into one digest. A shared tool holds a single API token with access to every client's CRM, and the agent picks the right account by name, which works right up until two clients have a contact called Sarah Chen.
Pick an Isolation Level Per Layer
You do not need one isolation strategy for the whole stack. You need a deliberate choice at each layer, written down, because the cheapest setting differs wildly between them. Here is how I would set a typical agency deployment serving five to twenty clients:
- Credentials: hard isolation, always. One vault path per tenant, one set of tokens per tenant, resolved at call time. There is no workload small enough to justify a shared token.
- Memory files: separate directories with a guarded root. The agent process gets a working directory scoped to one tenant and cannot resolve paths above it.
- Retrieval index: separate collections. Metadata filters on a shared index are tempting and cheaper, and they fail open the first time somebody writes a query without the filter.
- Agent process: shared, with per-request tenant binding. Running a separate gateway per client is cleaner and costs you an operations team. Most shops only need that at regulated clients.
- Logs and audit trail: shared storage, tenant-tagged, access-filtered. You want cross-tenant visibility here, for your operators only.
- Model provider: shared. Also the most dangerous layer, for reasons covered below.
Notice the pattern. Anything that holds client data at rest gets a hard wall. Anything that merely processes a request can be shared, as long as it carries the tenant identity with it the whole way down.
The Tenant Context Object
The single most useful piece of code in a multi-tenant OpenClaw stack is tiny. It resolves tenant identity once, at the gateway, from something the agent cannot influence (the inbound channel or the API key on the request), and then freezes it.
// resolved at ingress, never from model output
function resolveTenant(request) {
const tenantId = channelMap.lookup(request.source)
?? apiKeys.ownerOf(request.headers['x-api-key']);
if (!tenantId) throw new UnroutableRequest(request.source); // no default tenant
return Object.freeze({
tenantId,
workspaceRoot: `/var/openclaw/tenants/${tenantId}`,
vaultPath: `secret/tenants/${tenantId}`,
indexName: `mem_${tenantId}`,
budget: budgets.for(tenantId),
});
}
// every storage call takes the context as its first argument
memory.read(ctx, 'notes/renewals.md');
retrieval.search(ctx, 'renewal terms', { k: 8 });
tools.call(ctx, 'crm.find_contact', { name: 'Sarah Chen' });Two details in there carry most of the safety. There is no default tenant. A request that cannot be mapped to a client fails loudly instead of landing in whatever tenant was configured first, which is where every orphaned webhook in a sloppy system ends up. The tenant never comes from the model. If the agent can put tenantId in a tool argument, a prompt injection in one client's inbox can steer it into another client's CRM.
Then one more rule: storage functions refuse to run without a context, so forgetting to pass one is a crash in development instead of a leak in production. Sounds minor. It catches more bugs than everything else combined.
Memory and Retrieval
The memory architecture most OpenClaw setups use (flat files for durable facts, plus a vector index for recall) splits cleanly along tenant lines if you let it. Each tenant gets its own workspace root. The agent's file tools resolve every path relative to ctx.workspaceRoot and reject anything containing .. or an absolute path, before touching the filesystem.
Retrieval is where teams cut corners. A single shared collection with a tenant_id metadata field is one query builder bug away from returning everyone's data, and the bug is invisible because the results still look relevant. They just belong to somebody else. Separate collections per tenant cost a little more in index overhead and turn that bug into an error, because an unscoped query has no collection to run against.
Then there is the scratchpad. Long-running agents compact their context, and a compaction step that summarizes "recent work" across a shared process will happily blend two clients into one paragraph. Compaction has to run per tenant session, and the summary gets written back into that tenant's workspace only. The context window is the leakiest surface in the whole stack, because it is the one place where everything the agent knows is sitting in the same buffer at the same moment.
Credentials Resolve at Call Time
Environment variables are the classic single-tenant pattern and the classic multi-tenant failure. CRM_API_KEY can only hold one value. So people add ACME_CRM_API_KEY and GLOBEX_CRM_API_KEY, and now the agent has to pick the right variable, which means the agent's judgment is your access control.
Move the choice out of the agent. The tool layer receives the tenant context, fetches a short-lived token from ctx.vaultPath, makes the call, and discards the token. The agent never sees a credential and never names one. This is the same zero-trust posture the security architecture applies to a single deployment, extended one level: scope every token to the smallest owner that needs it, and at a multi-client shop the smallest owner is the client.
The Model Cannot Be Partitioned
Every other layer in this article can be walled off, but the model itself has no idea who it is working for, and the only thing standing between one client's secrets and another client's inbox on any given turn is whatever your code decided to put in the context window.
Noisy Neighbors and Per-Tenant Budgets
Leaks are the scary failure. Starvation is the common one. One client kicks off a backfill of eighteen months of support tickets, the job fans out to two hundred agent calls, and everyone else's morning digest arrives at noon because the queue is first-in, first-out and the backfill got in first.
Give every tenant its own queue and its own spend cap, then schedule across queues fairly. Weighted round-robin is plenty. A client on a larger plan gets a larger weight, a client in the middle of a backfill gets throttled to its share, and nobody's scheduled work waits behind somebody else's bulk job. The cost architecture patterns for model routing apply here per tenant, with one addition. Track spend by tenantId from the first day.
You will need that number when you price the next contract, and reconstructing it later from provider invoices is miserable.
Failure Modes Worth Knowing
The default tenant
A config fallback routes unmapped requests to the first tenant in the list. Months later a new client's webhook is misconfigured, and their events quietly land in another client's workspace.
Fix: delete the fallback. Unroutable requests go to a dead-letter queue that only operators can read, with an alert.
The omniscient operator agent
You build an internal admin agent that can see every tenant, for reporting. It works beautifully. Then someone wires it to answer a client's support question, and it answers with data drawn from the entire book of business.
Fix: cross-tenant agents never get a client-facing channel. If an operator agent must act for a client, it spawns a tenant-bound worker with a frozen context and hands the task down.
Prompt caching across tenants
Shared system prompts are great for cache hit rates. Shared system prompts with client-specific facts appended above the cache breakpoint put one client's details into the prefix that every other client reuses.
Fix: keep the cached prefix generic. Everything tenant-specific goes after the breakpoint.
Offboarding that forgets the index
A client leaves. You delete their workspace directory and revoke their tokens. Their embeddings stay in the vector store for a year, along with the summaries your nightly job wrote about them.
Fix: offboarding is a script that walks every layer from the list above using the same tenant context, and it runs in staging against a fake tenant before it ever runs for real.
Testing for Leaks on Purpose
Functional tests will never find a cross-tenant leak, because leaked data looks exactly like correct data. You have to plant it. Create two synthetic tenants, seed each with a canary string that exists nowhere else (a fake product code, a made-up employee name), and run your normal agent workloads against tenant A while grepping every output and every tool call argument for tenant B's canary.
Run that suite in CI on every change that touches retrieval or memory code. One canary hit fails the build. It is the cheapest insurance in the whole stack, and after the first time it catches something you will stop thinking of it as optional.
Internal Links & Further Reading
To go deeper on the layers this article references:
- OpenClaw Security Architecture: Authentication, Authorization, and Zero-Trust Patterns →
The credential scoping model that per-tenant vault paths extend.
- OpenClaw State Management →
Where session state lives, and how to key it so two clients never share a row.
- OpenClaw Human-in-the-Loop Architecture: Approval Gates and Autonomy Budgets →
Autonomy budgets work per tenant too, and some clients will want a stricter tier table than others.
- Monitoring and Observability: Seeing Inside the Agent Mind →
Tenant-tagged traces, per-client spend dashboards, and where to alert on canary hits.
FAQ
Q: At what point do I need any of this?
The day a second party's data enters the system. That includes a second department inside your own company if their data is confidential from the first. Retrofitting tenant context into a stack with six months of mixed memory files is far harder than adding it while there is only one tenant and the migration is a folder rename.
Q: Should each client get its own OpenClaw gateway instead?
For regulated clients (a hospital group, or anyone whose contract names data residency) a dedicated gateway and a dedicated machine is often the right answer, and the tenant context pattern makes that a deployment change rather than a rewrite. For everyone else, a shared gateway with hard walls at the storage and credential layers gives you nearly all the safety at a fraction of the operational load.
Q: Is a metadata filter on a shared vector index really that risky?
The filter itself works fine. The risk is the one code path that forgets to apply it, and in a growing codebase that path always shows up eventually, usually in a new feature written in a hurry. Separate collections make the unsafe query impossible to express. Filters make it merely incorrect.
Q: Can one agent persona serve multiple clients?
Yes. Persona and tenant are different axes. The same system prompt and skill set can serve every client, while the tenant context decides which memory and which credentials that persona can touch on a given request. Keep the persona in the shared, cacheable prefix and the tenant specifics after it.
The Bottom Line
Multi-tenancy in OpenClaw comes down to one discipline: resolve the client once, at ingress, from something the model cannot touch, then make every storage and credential call impossible to run without that identity. Wall off data at rest, share the processing, and keep a canary suite in CI that fails the build the moment one client's planted string shows up in another client's output.
Do it while you still have one client. The second one is where the pricing number leaks.
Skip the trial and error
Get the OpenClaw Starter Kit — config templates, 5 ready-made skills, deployment checklist. Everything you need to go from zero to running in under an hour.
$14 $6.99
Get the Starter Kit →Also in the OpenClaw store
Get the free OpenClaw deployment checklist
Production-ready setup steps. Nothing you don't need.