My OpenClaw Setup: One Box That Can Make Websites
- Category
- AI & Local LLM
- Posted
- July 31, 2026
- Updated
- September 16, 2026
- By
- Jacob Lloyd — written with AI assistance, post-project
- Read time
- 11 min read
In plain terms: This article explains how AI helpers on a small Linux computer in my house maintain this website. They write drafts, rebuild the site and prepare changes, but anything that goes live passes through one safety script that previews first, backs up the server and never deletes files. The helpers do the work; I decide when it goes live.
AI agents running on a small Linux box in my house maintain the website you’re reading. They can do real work on it, and the setup is built so they can’t wreck it: every change reaches the server through one script that previews first, backs up the server, never deletes a file, and publishes only when I say so. That script is the part worth copying, and most of this article is about it.
tl;dr
- What it is: a Linux mini PC running OpenClaw (a self-hosted AI agent gateway), wired to a static site generator and a gated deploy script.
- Where the thinking happens: the lead agent runs on a cloud model (MiniMax M3). Local models are served from a separate GPU machine on my network. The mini PC hosts the gateway, the site and the scripts, and runs no models itself.
- What it costs: it varies a lot by month; the real bills are in what my AI agents actually cost.
- What you end up with: agents that draft posts, rebuild the site and dry-run every deploy, behind a secret scan and a link check, with publishing kept as a human decision.
Status, September 2026: the weekly self-refresh described in Step 3 is off. It was lost in a restore from backup at the end of July and I haven’t recreated it. Everything else runs, on request.
The whole box at a glance
The gateway turns models into agents with tools. The agents edit plain-text content, a deterministic generator builds the site, and one gated script is the only road to the web server. The models themselves live elsewhere: a cloud API, and a GPU machine on the local network.
What you end up with
- “Draft me a post about X” in chat: an agent writes a Markdown article into the site’s content folder, following the frontmatter schema and house style, and rebuilds the local preview.
- “Deploy it”: the agent runs the deploy script, which by default is a dry run. It lists exactly which files would change on the server and touches nothing. Publishing for real needs an explicit
--gofrom me. - “Go over the site”: an agent sweeps for stale dates, weak posts and dead links, with the same guard checks in front of anything that would publish.
The agents do the typing, the scripts enforce the rules, and I make the one decision that matters: whether a change goes live.
The hardware: any Linux box works
Mine is a CHUWI AuBox Ai365 mini PC (the hardware details are in that article). Since the models run elsewhere, the website side needs very little:
- Any 64-bit Linux machine with systemd, for service supervision and timers.
- Python 3 for the static site generator, rsync + SSH for deploys.
- A model, either a cloud API key or a model server somewhere on your network. If you want everything on one box, LM Studio, Ollama or llama.cpp’s server all speak the OpenAI-compatible API a gateway expects.
Step 1: an agent gateway with one scoped agent
The gateway turns a model into an agent with tools, file access and scheduled jobs. I use OpenClaw: it hosts named agents, routes each to a model with fallbacks, and exposes a chat interface. Create one “webmaster” agent with file and shell tools, then apply two rules that hold for any gateway:
- Bind the gateway to loopback and put any chat UI behind authentication. An agent with shell access is not something to leave open on your network.
- Scope the agent’s working directory to the site repo. It needs to edit content, not your home directory.
My gateway hosts several other agents too (they’re described in the agent stack article). For the website, only one split matters: the family-safe bots that show up on a locked device have no file or shell tools at all, so a phone in the wrong hands can’t reach the site.
When the system itself breaks (a config regression, a crashed gateway, a build that won’t run), I escalate to a strong cloud coding model that treats the whole box as its patient. That story is in When OpenClaw Breaks, I Fix It with Claude Code.
Step 2: a static site plus a gated deploy script
Agents and WordPress mix badly: a database, a login, plugins and PHP are all attack surface and all failure modes. A static site suits an agent much better. The whole site is plain text files in a folder, the build is one command, and publishing is a file copy. My generator is under a thousand lines of Python (Jinja2 + Markdown + YAML); Hugo, Eleventy or Zola work the same way. What matters is the contract:
- Content is Markdown files with frontmatter. An LLM edits these natively, and a diff shows exactly what it did.
- The build is deterministic. Same content in, same site out, so any output change traces back to a content change. On my machine a full build takes about two seconds.
- There is exactly one deploy path, and it is a script with the safety built in, not a set of instructions the agent is trusted to follow.
That last one is the heart of the setup. The interface:
./deploy.sh preflight # checks tools, SSH, webroot, manifest, size budget
./deploy.sh # build + DRY-RUN rsync: prints what would change
./deploy.sh --go # build + server-side backup + publish + verify
What it guarantees:
- Dry-run by default. Running it with no arguments can’t modify the server, so an agent can run it freely to show me a pending deploy.
- Never deletes. rsync runs without
--delete. A confused agent can overwrite a page with a newer version, but it can’t erase the site. - Backs up first. With
--go, the script tars the server’s webroot before anything is overwritten and keeps the last five archives. If the backup fails, the deploy stops. - Fails closed. The deny rules come from a manifest file. A manifest that won’t parse, a manifest that yields no exclude rules, or a target that still holds its placeholder value all abort before rsync runs. Source files (scripts, templates, Markdown, YAML, keys) are also excluded by hard-coded rules, so they can’t ship even if the manifest is wrong. Posts marked as drafts are excluded too, and the script checks the transfer list to prove none slipped through.
- Verifies after. It fetches the live homepage and looks for a marker string the build embeds in every page. An HTTP 200 alone proves nothing if old hosting still answers for the domain.
Here is the core of it, trimmed and made generic. The real script is about 430 lines, most of them checks and error messages.
DRY=(-n); [ "$GO" -eq 1 ] && DRY=() # dry-run unless --go
# fail closed: no parseable manifest or no excludes = no rsync
m_raw="$(read_manifest)" || die "manifest parse failed"
mapfile -t M <<< "$m_raw"
[ "${#M[@]}" -gt 3 ] || die "manifest yielded no excludes"
MARKER="${M[2]}"; EXC=("${M[@]:3}")
HARD=(--exclude='.git/' --exclude='content/' --exclude='templates/'
--exclude='**/*.py' --exclude='**/*.sh' --exclude='**/*.md'
--exclude='**/*.yaml' --exclude='**/.env*' --exclude='**/*.key'
--exclude='**/*.bak*' --exclude='**/*~')
ssh "$DEST" "test -d $WEBROOT" || die "webroot not reachable"
if [ "$GO" -eq 1 ]; then # backup BEFORE overwrite
ssh "$DEST" "mkdir -p $WEBROOT-backups \
&& tar czf $WEBROOT-backups/site-$ts.tgz $WEBROOT \
&& test -s $WEBROOT-backups/site-$ts.tgz \
&& echo BACKUP_OK" | grep -q BACKUP_OK || die "backup failed, aborting"
fi
rsync -az --itemize-changes "${DRY[@]}" \
"${DRAFT_EXC[@]}" "${HARD[@]}" "${EXC[@]}" \
./ "$DEST:$WEBROOT/" # note: no --delete, ever
[ "$GO" -eq 1 ] || { echo "DRY-RUN only. Re-run with --go to publish."; exit 0; }
curl -s -L "$SITE_URL" | grep -qF "$MARKER" \
&& echo "marker found: the new site is live" \
|| echo "HTTP answered but marker missing: old host still serving?"
On top of that sits a one-word wrapper, publish, which chains the steps in order and stops at the first failure: build with a broken-link and missing-image check, a local backup of the source tree, a pre-publish scan of the built HTML, then the deploy script with --go.
The agent may run preflight and the dry run whenever it likes. The --go flag is reserved for me. That one asymmetry is what makes it reasonable for an AI to sit this close to a production site.
Step 3: putting the refresh on a timer (and why mine is off)
The last piece makes a site self-maintaining, and it’s the piece I currently run by hand. A systemd timer fires on a schedule and hands the agent a standing brief: review the site, refresh anything stale, tighten weak posts, check internal links, and stage the result without publishing it.
# ~/.config/systemd/user/site-refresh.timer
[Unit]
Description=Weekly website content refresh
[Timer]
OnCalendar=Sun 06:00
Persistent=true
[Install]
WantedBy=timers.target
The service it triggers calls the gateway’s CLI or API with the brief, then runs the guard checks on whatever the agent staged:
- Secret scan. A pattern scan for anything shaped like an API key, token, password, private hostname or private IP. gitleaks does this off the shelf. Any hit aborts the run.
- Content check. Mine is a word-list scan for adult or unprofessional language. A second model reading the draft with a different prompt catches things a word list can’t, like claims that aren’t backed by anything.
- Build must pass, including the link and asset checker, so a broken image or dead internal link fails the run before it reaches the deploy script.
Only when all three pass does the job dry-run a deploy, and publishing still waits for my --go. If you ever let a scheduled job publish on its own, the guards and the backup-first, never-delete script cap the damage at “a bad article went live and I restored from backup.”
Mine ran this way until a restore from backup at the end of July quietly took the job with it. I noticed weeks later, from a date on this page that had stopped moving. That is the failure mode of every scheduled job: a timer that stops firing looks exactly like a timer with nothing to do. If you build this, make the job report something every run, even “nothing changed”, so that silence means broken.
What the guards have actually caught
Two real cases from this site, both worth designing for:
- A scanner that skipped a file type. My secret scanner skipped
.zipfiles, and the site’s content scan never looked in the downloads folder. Three published download archives turned out to contain private details. The fix was to make both scanners open archives and scan the files inside, with size limits so a hostile archive can’t hang the check. The question to ask of any scanner is what it declines to open. - An agent draft that invented a fact. A draft about one model vendor stated, confidently, that its model was sold under a competitor’s product name. It wasn’t, and the sentence got through four rounds of review. Now every number and vendor claim in a draft needs a source in a separate facts file, and a claim with no source gets cut.
Replicating this
Nothing here is proprietary. The recipe:
- Pick a Linux box and decide where the model comes from: a cloud API, a local model server, or a GPU machine on your network.
- Install an agent gateway (OpenClaw or similar) bound to loopback, with one webmaster agent whose file and shell tools are scoped to the site directory.
- Move the site to a static generator if it isn’t on one already.
- Write the gated deploy script with the five guarantees above. This is the piece worth the most care. Hand the guarantee list and the excerpt to your own LLM as a spec.
- Add the guard checks (secret scan, content check, link check), all fail-closed, and a timer if you want one, with a report every run.
- Keep
--gohuman until the pipeline has earned otherwise.
Gotchas
- The deploy script must be the only path. The first hand-rolled rsync “just this once” is the incident the script exists to prevent. Mine hard-excludes source files and refuses placeholder configs because a one-off command wouldn’t.
- A bare HTTP 200 is not a successful deploy. During my WordPress migration the old hosting kept answering for the domain, so every check passed while the new files sat unused. Embed a marker in your build and look for it on the live page. The reverse also happens: the host can serve a cached homepage for a minute after a good deploy, so re-check once before calling a missing marker a failure.
- Agents copy the schema you show them, placeholders included. If your post scaffold contains placeholder text, a lazy model will publish the placeholder excerpt word for word. Add your scaffold’s placeholder strings to the content check.
- Don’t let the writer review itself. Same model and same prompt means the same blind spots. Give the reviewer a different prompt at minimum, and ideally a different model.
- Scheduled LLM jobs need timeouts and a “staged, not published” default. A job that can publish unsupervised will eventually publish something odd at 6am on a Sunday. Staging costs you one review.
- Secret scans belong in the pipeline, not in the agent’s instructions. “Never include secrets” in a prompt is a wish. A fail-closed scan on the output is a control.
Related: my agent stack, what the agents actually cost, when OpenClaw breaks, the mini PC it runs on, building a second site with the same agents and the WordPress days before this.