# A Small Office of AI Agents on One Mini PC: My OpenClaw Stack URL: https://www.laserlloyd.com/projects/local-ai-agent-stack-openclaw/ Published: 2026-07-11 Updated: 2026-09-16 Description: How I run six AI agents from one AMD mini PC with OpenClaw: a lead, a stateless worker, a reviewer, and local models on a GPU box, with config shown. Reuse: free for personal use — full policy at https://www.laserlloyd.com/llms.txt Six AI agents run from a mini PC in my house. One talks to me and decides who does what. One is a worker that gets spawned fresh for every heavy job. One checks finished work. The rest are persona agents, plus one safe agent the family can open from a phone. The lead runs on a flat-rate cloud plan and the private jobs run on local models, so the monthly bill depends mostly on how much metered work I push through. This article shows how the pieces fit, the config that wires them, one real hand-off from start to finish, and the watchdog that makes sure finished work reaches me. tl;dr What it is: OpenClaw, a self-hosted AI agent gateway, running six agents, each with its own job, model fallback chain and tool permissions, behind a chat app I built. The shape: a lead that triages, one stateless worker with four model variants, a reviewer on a different model, and local models served from a GPU machine on my network. What it costs: it varies a lot by month. The real bills are in what my AI agents actually cost; this article covers what drives them. What you need: a Linux box with 32GB of RAM (mine is a CHUWI AuBox mini PC) and some patience for systemd. A second machine with real GPUs is optional and changes what the local tier can do. What you end up with One chat URL for the whole house. You ask the lead for something, it either answers or hands the job to a worker, and the result comes back into the same thread. Family devices see only the safe agent. Image generation happens on the GPU box, and email drafts wait for a human to send them. The chat front end is DisPatch; these shots are sandbox demo data from that write-up, from July 2026. Desktop, unlocked: agent rail on the left, the lead agent in the chat The same chat on a phone Locked mode, what family devices see. The July 2026 demo had three safe bots; since the September 2026 rebuild there is one. The box The gateway, the chat app and the scripts run on one small machine: CHUWI AuBox Ai365: AMD Ryzen AI 9 365, 10 cores, Radeon 880M integrated GPU. 30GB of usable DDR5, shared between the OS and everything else, so there is no real VRAM to speak of. Bluefin, an immutable Fedora Atomic system. It's an AMD machine, so anything GPU-flavoured here means ROCm, not CUDA. It is the second box to hold the job. The first was a 128GB ROG Flow Z13 that ran a 120B mixture-of-experts model locally until it developed a power fault in July. The replacement has a quarter of the memory, so the big models moved to a second PC on my network with real GPUs, served by StudioForge, an OpenAI-compatible server I wrote for that job. Image generation moved there too. Local ComfyUI used to run on this box and no longer does. The mini PC runs no models at all now. It used to keep LM Studio for a 7B vision model and an embedding model; I removed LM Studio from it on 2026-09-16, and the GPU PC covers those jobs. The cast The fleet was rebuilt on 2026-09-06. It had grown to a dozen agents with overlapping jobs, and it came out as six. The rule I kept: one agent per kind of decision, and every agent gets only the tools that job needs. RoleJobModel (fallbacks)Tools LeadTalks to me, answers small things, runs quick shell commands, spawns workers for anything heavyMiniMax M3 on a flat-rate plan (DeepSeek pro, then flash)Files, shell, web search, spawning. No browser. WorkerStateless. A fresh session per job, then it's gone. Cannot spawn anything itselfMiniMax M3 by default (DeepSeek flash, then pro). Four variants, belowFull: files, shell, web, the GPU box's model and image tools, mail drafting ReviewerSpot-checks finished work and gates anything that leaves the houseKimi K3 (DeepSeek pro)Files, shell, browser, session history CompanionA persona agent for long conversations; pictures come from a fixed script, not a model callA 27B model on the GPU box (MiniMax M3)Files, shell, web search CoachPersonal-trainer personaThe same 27B on the GPU box (MiniMax M3)Files, shell, web search, reading images Safe agentThe only agent a locked family device can reachA fast MiniMax model (DeepSeek flash, then pro)Read and web search only. No shell, no image tools, no files outside its own folder. The worker's four variants are model aliases the lead picks when it spawns a job: fast: DeepSeek flash, for lookups and routine edits. think: MiniMax M3, the default for real work. local: the 27B on the GPU box, for private or bulk jobs that shouldn't touch a cloud API. deep: MiniMax M3 through a second provider route with a 1M-token window, for long multi-file jobs. Having one worker with four variants, instead of four named workers, removed a whole class of mistakes. Every worker used to carry its own long-lived session and its own copy of the rules. Now the rules live in one place and every job starts clean. Since the rebuild I've added a few specialist coaching seats and bridged in two external coding agents, but the six above do the daily work. What the config looks like Everything above lives in ~/.openclaw/openclaw.json. This is a trimmed version of my agent block with the ids renamed to their roles. The keys and values are the real ones from OpenClaw 2026.9.4. { "agents": { "defaults": { "models": { "deepseek/deepseek-v4-flash": { "alias": "worker-fast" }, "minimax/MiniMax-M3": { "alias": "worker-think" } }, "subagents": { "maxSpawnDepth": 2, "maxConcurrent": 8 } }, "entries": { "lead": { "model": { "primary": "minimax/MiniMax-M3", "fallbacks": ["deepseek/deepseek-v4-pro", "deepseek/deepseek-v4-flash"] }, "subagents": { "allowAgents": ["worker", "reviewer"] }, "tools": { "profile": "minimal", "alsoAllow": ["sessions_spawn", "read", "write", "edit", "exec", "web_fetch"], "deny": ["browser"] } }, "worker": { "model": { "primary": "minimax/MiniMax-M3", "fallbacks": ["deepseek/deepseek-v4-flash", "deepseek/deepseek-v4-pro"] }, "subagents": { "allowAgents": [] } }, "safe": { "tools": { "allow": ["read", "web_fetch", "brave_search"], "deny": ["exec", "process"], "fs": { "workspaceOnly": true } } } } }, "tools": { "agentToAgent": { "enabled": true, "allow": ["lead", "worker", "reviewer", "safe"] } } } Three things in there do most of the work: fallbacks is an ordered list. If the primary is down or rate-limited, the gateway tries the next one, so a cloud outage turns into a slower answer instead of no answer. "profile": "minimal" plus alsoAllow is a strict allowlist: the agent gets the minimal set and exactly the tools you name. deny beats both, so a denied tool can't be granted back by accident. "allowAgents": [] on the worker makes it a leaf. It can't spawn helpers of its own, so one job can't fan out into ten. One trap: a worker spawned by the lead can only use tools the lead also holds. On the gateway version I rebuilt on, the spawned worker came up with seven tools and no shell until the lead was given exec too. If a worker seems oddly limited, compare its tool list with its parent's. Handing off real work Chat replies are the easy case. The useful pattern is asking for something big and letting the lead decide who does it. 1. You ask the lead agent "fix this launcher," "audit the site" 2. The lead triages small things itself (quick shell), big things dispatched 3. A worker session is spawned with a written brief and a model variant → 4. The worker runs the job files, shell, web, GPU-box tools, in its own session 5. The result reaches your thread via the lead, with a scanner as backstop Picked for the job fast: DeepSeek flash think: MiniMax M3 (default) local: 27B on the GPU box deep: M3, 1M-token route or the reviewer (Kimi K3) for checks the ask spawn runs report written picks The lead is allowed to do small things itself. For a while it was barred from running any shell command at all, which broke simple jobs like posting its own morning summary, so its shell access came back. The working rule is in its instructions: a couple of quick commands are fine, and anything longer goes to a worker. A real example On 2026-09-15 I told the lead that the desktop shortcut for a coding agent's web UI had stopped working. The UI now issues a fresh access token every time its service restarts, and the shortcut pointed at the bare URL, so it got a 401. The lead didn't fix it itself. It wrote a brief and spawned a worker. This is the brief's shape, shortened: Task: Update the web-UI app shortcut so it survives the per-restart token rotation. Context: The UI serves only with a rotating ?token= query; the launcher points at the bare URL and gets 401. Sources: the .desktop file; the tool's --help; its source, for any flag or env var that gives a stable token, before reaching for a wrapper. Constraints: No service restart, no config edits. Write surface: one .desktop entry and at most one wrapper script (~30 lines). Do NOT bake a static token into any file. Output: Result / Evidence / Files / Failed / Next. Stop rules: Two failures of the same tool = stop and report. The worker read the tool's help and its source and found no stable-token option: the token is random per process and never read from the environment. So it wrote a 20-line wrapper that pulls the current token from the service's own startup line in the journal and opens that URL, falling back to the plain URL if the line is gone. Then it pointed the shortcut at the wrapper. The run took 125 seconds and 46 tool calls. It ran on the worker's default model, which is on a flat-rate plan, so the job added nothing to a metered bill unless a fallback kicked in. Three parts of that brief do the real work. The write surface is named, so the worker can't wander off and edit config. The obvious wrong fix (a hard-coded token) is ruled out in writing. And the stop rule keeps a stuck worker from looping on the same failing command. Making sure finished work arrives A dispatched job that silently dies is worse than one that never started, because you don't find out until you go looking. I've built this check twice. Version 1: a path unit and a five-minute timer (June to September 2026) 1. Dispatch logged the lead appends a row to a tracker file 2. systemd path unit fires watches the file, triggers on change 3. A 5-minute timer arms one-shot, self-deleting, one per dispatch Worker reported in time the check exits quietly, timer is gone Worker stayed silent a worker agent is sent to chase it file changed runs the arming script report arrives 5 minutes pass It replaced a bad first attempt: a cron job that woke an LLM every five minutes to read the tracker. That spun a model on every tick and caused the very request timeouts it was supposed to catch. The event-driven version costs nothing until there is a dispatch to watch. The path unit is the copyable part: # ~/.config/systemd/user/dispatch-watch.path [Path] PathModified=%h/agents/dispatch-tracker.md Unit=dispatch-watch.service [Install] WantedBy=default.target # ~/.config/systemd/user/dispatch-watch.service [Service] Type=oneshot ExecStart=%h/bin/dispatch-arm.sh And this is the core of the arming script, condensed. It asks a small Python helper for the keys of still-open dispatches, then arms one transient timer per new key: for key in $(dispatch-watchdog --list-open); do grep -qx "$key" "$ARMED" && continue # one follow-up per dispatch, ever systemd-run --user --on-active=5min \ --unit="dispatch-followup-$key" \ dispatch-watchdog --check "$key" && echo "$key" >> "$ARMED" done When a check found a dispatch still open, it started a worker agent with a triage prompt. The worker looked at the session list and answered with one codeword: DONE, STILL-RUNNING, FIXED or ERROR. The Python helper then wrote the status into the tracker, so two agents never edited the same file. A hung job got re-dispatched once, and I got a ping. Version 2: run folders and a standing scanner (since 2026-09-06) Version 1 had a blind spot. It only knew about jobs the lead wrote down, and it assumed the lead would still be listening when the worker finished. Sometimes it wasn't. A worker can finish after the session that spawned it has ended, and the gateway then retries its completion notice for about 30 minutes before giving up. The work was done and nobody heard about it. Now every worker's last act is to write two files into its own run folder: report.md and a small meta.json. A timer that never stops walks those folders and handles anything the normal path missed (paths and flags trimmed): # runs-deliver.timer [Timer] OnBootSec=2min OnUnitActiveSec=45s AccuracySec=5s # runs-deliver.service [Service] Type=oneshot ExecStart=%h/bin/runs-deliver --once --root %h/agents/runs Nice=10 IOSchedulingClass=idle TimeoutStartSec=40 { "id": "run-20260915-125334", "kind": "spawn", "status": "done", "created": "2026-09-15T12:53:34Z", "finished": "2026-09-15T12:55:31Z", "delivered": true } A successful run gets no chat message of its own, because a daily summary carries the count. A failed run gets one line in the thread. A run that finishes more than three hours after it started is reported as recovered. The example above went back to the lead through the gateway's normal path, and the scanner marked it handled 13 seconds after the worker finished. The lesson from both versions: make the worker leave a file behind, and have something with no stake in the conversation read it. What drives the cost I keep the dollar figures in one place, the cost article, because they move month to month and provider to provider. What decides the bill is where the volume goes: Flat-rate lead and worker. The model that does most of the talking and most of the work is on a subscription. Heavy use doesn't grow that bill; it just runs into the plan's limits. Metered fallbacks. The DeepSeek models behind them are billed per token. A day with provider trouble on the primary quietly shifts traffic onto the meter. Local models cost electricity. The companions, the local worker variant, and memory search (a local embedding model) never touch a paid API, so the constant background lookups are free. Long sessions cost the most. Every turn re-sends the conversation so far. Capping context and compacting long sessions is the first lever I reach for when a bill jumps. The rules it won't break A gateway that can spawn agents and run tools needs guardrails that hold even when the model misbehaves: Camera, screen recording, SMS, contacts, calendar and reminders commands are on the gateway's deny list. The agents can't call them at all. The safe agent has an explicit allowlist of read and web-search tools, with shell and image tools denied on top of that. Plugins are a strict allowlist of eleven entries: five model providers, web search, a headless browser, two memory plugins, the email bridge and a bridge for external coding agents. Nothing else loads, installed or not. Anything that publishes (site deploys, email sends) dry-runs by default, and the real run needs my explicit go. Email drafts are never sent by an agent. Gotchas These actually broke, roughly in the order they bit me: Local reasoning models need reasoning: true set explicitly. Mine send a separate reasoning stream before the answer. With the flag missing, the gateway read the silence as a stall and killed the request around 6.5 minutes in. Every timeout defaults to cloud speed. A big model on slow hardware looks stalled when it's just thinking. My provider request timeout is 20 minutes and the agent turn timeout is an hour. Before I raised them, long local jobs got killed partway through. LM Studio's AppImage crashed with SIGBUS (back when this box still ran it) whenever its /tmp FUSE mount got recycled under memory pressure. It was a real bus error, and the OOM killer had nothing to do with it. Extracting the AppImage and running it as a plain systemd service fixed it. The messaging allowlist is case-sensitive, but the gateway lowercases agent IDs. Cross-agent sends failed with no error and no log line until I used lowercase IDs everywhere. Cross-agent messaging once needed two switches: the feature flag and the session-visibility setting. That was true of the mid-2026 gateway; newer releases default visibility to open, so check your version before debugging. A shell redirect ate my watchdog's errors. The arming script opened its lock file with exec 9>"$LOCK" 2>/dev/null. With no command, exec applies every redirect to the whole shell, so all later error output vanished, including "failed to arm". Guard only the lock creation, then exec 9>>"$LOCK". Reusing one session for every escalation made it rot. Version 1 sent every chase into the same worker session. It grew to 5.75 MB and started stalling. Each escalation now gets its own session key. An agent once resurrected its own bad config. An identity-reconcile pass re-injected a stale workspace file and crash-looped the gateway. Fix the source file, not only the config. A non-empty plugin allowlist is absolute. enabled: true does nothing if the plugin's ID isn't also in the allow list, and nothing tells you. The same trap is in the model manager write-up. Related: how the same box maintains this website; MiniMax M3 as a coding agent on OpenClaw; and the benchmark I use to pick models.