Home › AI & Local LLM
AI & Local LLM
Running AI on your own hardware — local LLMs, agents, and image generation. No subscriptions, no cloud, no computer-science degree.
What My AI Agents Actually Cost: DeepSeek vs the Big APIs
Three months of an always-on AI agent stack: $242 in June, $225 in July, $385 in August. What drove each bill, what DeepSeek's August 16 price hike did, and what the same workload costs on GPT, Claude, Grok and Kimi.
StudioForge: A GPU-Only LLM Server That Replaced LM Studio
I wrote the model server I wanted: an OpenAI-compatible llama.cpp gateway on port 1234 that plans VRAM before it loads, refuses a model instead of spilling it to the CPU, keeps pinned models warm, and gives an agent on another machine 29 MCP tools to run the rig. Free download, source on GitHub.
MailForge: A Local-LLM Email Assistant That Never Sends
A self-hosted email triage assistant: an IMAP listener, a local model, and a drafts queue. It classifies, labels and writes the reply, then stops. Sending is a human click, always. Free, MIT, download below.
DisPatch: A Self-Hosted Chat App for Local AI Agents
A self-hosted chat app for local AI agents, with a drag-drop file server and a server-enforced PIN lock. It replaced my 30-container Rocket.Chat install. Version 2.0 is out, and the download is a 3 MB zip.
ThemeForge: a Drop-in CSS Theme System an AI Agent Installs in One Step
ThemeForge is a single folder, `ui-theme/`, that gives a web app ten colour themes, a picker, design tokens and ready-made component classes — and it's written so an AI coding agent can install it in one step from a one-line Python command. Plain CSS, no build step, no dependencies; MIT; source and zip below.
ChatForge: a Local NPU AI Assistant for the Copilot Key
My laptop shipped with an AI key and an AI chip that had never met. ChatForge introduces them: press the Copilot key, get a 1.5B Qwen answering on the Intel NPU at 42–51 tokens/s — no account, nothing leaves the PC — with StudioForge or a cloud model one click away. MIT, source and zip below.
Winding Down OMLA: How to Shut Down a Non-Profit
I spent months building a non-profit licensing framework for open-weight AI models, then retired it without ever launching. What it was, what got built, and what shutting it down actually took — with the whole thing archived for download.
WordPress: Use AI to Make Custom Blocks (Free Accordion Example)
A free copy-paste accordion for WordPress — plain HTML and CSS, no plugin, no JavaScript — plus the exact AI prompt that wrote it.
Putting a Chatbot on the Public Web — On My Own Hardware
A chatbot on a public website, but the brain is a small model running on a computer in my house — free per token and private. Here's the relay architecture that makes it safe, the security holes I had to close, and an honest accounting of how slow it is and what it would still need to be production-ready.
Set Up Your LLM Assistant: Claude Code, Cowork, Reasonix and DeepSeek
The hub guide every "implement this yourself" box on this site points at: what an agentic LLM assistant is, and four real ways to get one running today.
Reasonix: A Claude-Code-Style Coding Agent on DeepSeek
A terminal coding agent that works like Claude Code — subagents, skills, project memory — but runs against DeepSeek's pay-per-token API instead of a subscription. Setup, model tiers, and a hands-off auto-updater.
Pixel-Art Widget Showdown: Three Coding Agents, One Brief, Two Authors
I gave dsh, Kimi K3 and MiniMax M3 the same brief for a pixel-art LaserLloyd logo widget. dsh and Kimi wrote working widgets. MiniMax replied that it had built one, but the files it pointed to were Kimi's.
OpenClaw Broken? I Fix It with Claude Code
My two-tier repair strategy: OpenClaw agents handle the routine work, and when something truly breaks, Claude Code fixes the box — then teaches OpenClaw to handle it next time.
MiniMax M3: An OpenClaw Coding Agent vs Kimi K3 and dsh
MiniMax M3 on OpenClaw's own agent loop passed five small coding tasks in 159 s against Kimi K3's 310 s (Kimi at max thinking, different fixtures), at a quarter of the list-price cost. The widget I first credited to it was Kimi's.
My OpenClaw Setup: One Box That Can Make Websites
A Linux mini PC runs OpenClaw agents that write, build and publish this website. The one road to the server is a deploy script that dry-runs by default, backs up first, never deletes, and waits for my go-ahead.
Local AI Image Generation on AMD (ROCm): ComfyUI + Z-Image Turbo
How I turned an AMD Strix Halo machine into a local AI image generator — ComfyUI and Z-Image Turbo running as a systemd service, no cloud account, about 27 seconds per 1024px image.
A Small Office of AI Agents on One Mini PC: My OpenClaw Stack
Six AI agents run from a mini PC in my house: a lead on a flat-rate cloud plan, one stateless worker that does the heavy lifting, a reviewer, and local models on a GPU box. Here is the config, a real hand-off, and the watchdog that keeps finished work from vanishing.
Kimi K3 in OpenClaw vs dsh: 5/5 Tasks, 9× Slower, 48× the Cost
Moonshot's Kimi K3 driving OpenClaw's agent loop on five small coding tasks: 5/5 passes, about 9× slower and 48× more expensive than dsh on DeepSeek V4-Flash's warm pass (22× against its cold pass). The pixel-art widget brief took three attempts.
InfoForge: Offline Wikipedia as a RAG Source for My Local AI Agents
All of English Wikipedia — 19.19 million entries, 52.69 GB, no internet — sitting behind a local search index my AI agents can query. Every answer cites its articles, and when it can't find one it says so instead of guessing.
How I Built OMLA by Directing AI Agents Instead of Writing the Code
One guy, no backend team: I directed AI coding agents to build a real royalty-licensing platform for a nonprofit, audits and gated deploy included. The actual process, not the highlight reel.
How to Have an LLM Adapt Any Project to Your System
The general recipe for taking any open project — a GitHub repo or a zip from this site — and having an AI assistant make it run on your machine, even if you don't code.
Headless ComfyUI: Run It as a Service, Use It from Anywhere
The follow-up to my AMD ComfyUI build: run it headless as a systemd service, drive it from scripts over the HTTP API, and use the browser UI securely from any device — while every workflow still saves to the box.
I Built My Family a Japan Trip Itinerary App With AI Inside
A private, PIN-gated itinerary web app for a family trip to Japan — day-by-day plans, live weather, offline support, and a built-in AI assistant that can answer questions and make small edits. Here's how it came together, gotchas and all, with a scrubbed sample you can download.

DeepSeek: Running Locally — a 4-Step Guide (No Experience Needed)
A four-step guide to running DeepSeek on your own PC with Ollama, Docker, and Open WebUI. No cloud account, no subscription — just a graphics card with 8 GB of vRAM.
DeepSeek Harness (dsh): Install, Headless Runs, 3-Model Benchmark
DeepSeek's MIT-licensed coding agent, installed five days after release: npm install, a headless one-shot mode built for scripts, a systemd unit, model switching in one YAML file, and 15/15 passes on V4-Flash, V4-Pro and a local Gemma for a few cents.
DeepSeek Everywhere: Claude Code and a Local Agent Stack
In July 2026 DeepSeek quietly backed both my home agent stack and Claude Code. Here are the three copy-pasteable wiring patterns, including the systemd trick that keeps the API key out of argv entirely.
The CHUWI AuBox Ai365: My AI Home Lab in a Box
An obscure Strix Point mini PC with dual 2.5GbE, USB4 and a Ryzen AI 9 365. It runs Bluefin Linux on in-tree drivers, hosts my six-agent AI stack and draws 4.6 W at the chip. It replaced a dead 128 GB tablet.
My AI Agent Built a Benchmarking Tool for My AI Agents
My coding agent wrote bench-llm: 1,660 lines of Python that benchmark local models in LM Studio. Within two weeks my models all scored 5/5, so I replaced it with CrucibleForge: 251 cases, 158 hard, now on GitHub.
ASUS ROG Flow Z13: My Portable Home Lab That Refused to Stay Portable
A 128 GB gaming tablet ran my AI agents, my websites, and my whole home network — until a power fault started dropping it into sleep 30 seconds after every power-on. Here's the full post-mortem.
WordPress to Static With Claude Code: A Site Rebuilt in an Afternoon
This site used to run on WordPress. An AI coding agent rebuilt it as Markdown files plus a 183-line Python generator in an afternoon — and the same workflow still maintains it.

AI Chef: A Free System Prompt for Weekly Meal Planning
One copy-paste system prompt turns a chat AI into ProChef: weekly dinner plans, real recipes, and a grocery list already sorted by store section. Runs on my local Ollama box, but plain ChatGPT works too.