# StudioForge: A GPU-Only LLM Server That Replaced LM Studio URL: https://www.laserlloyd.com/projects/studioforge-gpu-only-llm-server/ Published: 2026-08-23 Updated: 2026-09-16 Description: StudioForge: a GPU-only llama.cpp server on port 1234 with a VRAM planner, model pins, GPU leases and 29 MCP tools. Free, MIT. Reuse: free for personal use — full policy at https://www.laserlloyd.com/llms.txt StudioForge is the local model server I wrote to replace LM Studio on my GPU rig. It supervises llama.cpp, speaks the OpenAI API on port 1234 (LM Studio's port, on purpose), has a browser control panel and a recovery sidecar, and publishes its management plane over MCP so an agent on another machine can run it without a shell. Full source on GitHub or download the zip I built it after too many LM Studio surprises: a 200 for unrouted paths, no way to ask what was actually loaded, and reasoning models overflowing an 8,192-token window my client config never asked for. The costly one was a model that did not fit being cut down instead of refused. LM Studio documents that it "will automatically reduce the GPU offload size … and the rest in the system RAM", which runs at system-RAM speed and still reports success. If you read the dsh write-up, the provider block called "Local GPU rig" is this server. It is also what my local agent stack, DisPatch, MailForge and InfoForge call for models. tl;dr What it is: a GPU-only, OpenAI-compatible server around llama.cpp's llama-server. Version 0.2.0, measured on llama.cpp build b10425 (CUDA 13.3). Gateway on 1234, control panel on 8080, watchdog on 1235, one child process per loaded model on 18100–18200. What it does: loads a model on first use, choosing context, KV cache type, GPU placement and slot count against the VRAM free at that moment. It idles models out, keeps pinned ones loaded, hands whole cards to one model on request, and publishes 29 MCP tools in 0.2.0 (19 management, 10 recovery). What it never does: spill to the CPU. A model fits entirely in VRAM or is refused with the numbers. It makes no outbound calls except Hugging Face for models, GitHub for the pinned engine and its update check, an opt-in release check, and image URLs a request names. What you need: NVIDIA GPUs on a 580-series-or-newer driver, Python 3.12+, uv and a folder of GGUFs. Never run a model locally? Start here instead. Limits: NVIDIA only. Windows is the reference platform. Linux is supported and tested in CI, but I have not run the Linux engine source build end to end myself. MIT-licensed, with one LGPL dependency covered under Limits. What it is not: Not a model or an engine. llama.cpp does the maths. StudioForge decides which process runs with which flags on which cards. Not a chat app. The Chat tab exists to prove the real request path works. Not a CPU server. --n-gpu-layers is 999 everywhere in the codebase. Not a second API. No Ollama or KoboldCpp endpoints. You get the OpenAI surface, an LM Studio-style /api/v0 mirror and the /api management REST. How it compares LM Studio 0.4.21StudioForge 0.2.0 Engineown llama.cpp builds plus MLX on Appleupstream llama-server, one pinned build (b10425, CUDA 13.3), smoke-tested before use Unrouted paths200 with an error body; its log says Returning 200 anyway404 with a JSON envelope, JSON on every status Errorsprose that clients regex-matcha stable error.code, diagnostics under error.studioforge Load configcontext_length ignored on one of two load paths; repetition_penalty silently ignoredone load path, every field honoured, effective values echoed back When it does not fit"will automatically reduce the GPU offload size … and the rest in the system RAM"507 insufficient_vram with required and available bytes, free VRAM per GPU, the largest context that would fit, and suggestions Multi-GPUpriority or even split, per-GPU toggles, tensor parallel since 0.4.15a planner that sizes context, KV type and slots per placement Idle TTL60 minutes; auto-evict keeps at most 1 JIT-loaded model1,800 s (30 min) shipped, 900 s (15 min) on my rig, swept every 15 s; as many models as fit Remote management/api/v1 load/unload/download; LM Link in preview/api REST, the sfctl CLI, 29 MCP tools, over your LAN or your own mesh VPN MCPhost only (it uses MCP servers)an MCP server for its own management, plus one on the watchdog Sourceclosed; free for personal and internal business useopen, MIT LM Studio rows checked against its changelog, docs and bug tracker on 2026-08-23 (version 0.4.21, released 2026-08-12). Row 2 is in its public bug tracker. Rows 3 and 4 are what my own client had to work around on the 0.3.x API; I have not re-tested them on 0.4.21, so treat them as documented history. LM Studio is the better desktop app: a polished GUI, MLX on Apple Silicon, an Anthropic-compatible endpoint, a mobile companion and frequent releases. The real difference is the default. LM Studio tries to make a model run somehow. StudioForge refuses and tells you why. When an agent decides what loads, that matters, because nothing downstream can tell a fast load from a crippled one without measuring. Server · engine · licenceHot-swap / idle TTLMulti-GPU placementRefuses CPU spillRemote mgmt / MCP StudioForge 0.2.0llama.cpp, one pinned build · MIT✅ JIT · 1,800 s TTL, 15 s sweepplanned per model, mixed cards, pins + leases✅ -ngl 999 + --fit off, 507 with numbers✅ REST + 29 MCP tools LM Studio 0.4.21own llama.cpp + MLX · closed✅ JIT · 60 min, auto-evict to 1priority/even split, tensor parallel❌ reduces offload, rest in RAMREST; MCP host only Ollama 0.32.15llama.cpp/GGML; MLX on Apple · MIT✅ keep_alive 5 min · 3 per GPU residentauto-spread across cards❌ spills, shows CPU % in ollama psrich /api/*; no MCP server llama-server router (b105xx, August 2026)is llama.cpp · MIT✅ child per model · --sleep-idle-seconds, LRU by countmanual -sm / -ts / -dev❌ --fit on shrinks your plan/models/load|unload llama-swap v251a proxy that spawns others · MIT✅ the whole product · per-model/group ttl❌ whatever your cmd saysn/a (proxy only)/ui + upstream routes vLLM 0.27.1own (PagedAttention) · Apache-2.0❌ one model per processtensor / pipeline / expert parallelpartial (no layer-offload path)LoRA only, "local dev" KoboldCpp 1.119llama.cpp fork + image/audio · AGPL-3.0✅ --admin + --routermodemanual --tensor_split❌ spills/api/admin/*; MCP client only TextGen (ex-oobabooga) 4.95 loaders incl. ExLlamaV3, TRT-LLM · AGPL-3.0✅ switch without restart · TTL ?manual --tensor-split❌ spills/v1/internal/model/*; MCP client TabbyAPI (rolling)ExLlamaV3 only, no GGUF · AGPL-3.0✅ admin + inline loading · TTL ?gpu_split_auto on by default✅ in effect (ExLlama has no CPU path)admin-key /v1/model/load Jan 0.8.4llama.cpp router mode · Apache-2.0✅ via the router · TTL ?inherited from llama.cpp❌ spills/v1/orchestrations; MCP client LocalAI 4.9.060+ backends as container images · MIT✅ on-demand · WATCHDOG_IDLE_TIMEOUT"automatic GPU model fitting"❌ "No GPU required"REST + UI; MCP client only GPUStack 2.2.3vLLM, SGLang, MindIE, VoxBox · Apache-2.0partial (cluster deployments)auto Spread/Binpack, multi-node?full cluster management API Checked from primary sources on 2026-08-23. A question mark means unknown, not "no". GPUStack's workers are Linux-only; vLLM has no native Windows support. JIT loading and idle TTL are not new. LM Studio, Ollama, llama-swap and LocalAI all do both, and llama-swap's per-group swap / exclusive / persistent flags are a neat policy engine. Three things are rare: refusing to spill to CPU (only TabbyAPI matches, because ExLlama has no CPU path), a planner that sizes context and slots to the cards it found, and management published as MCP tools, which I could not find anywhere else in this category. The closest rival is llama.cpp's own router. It runs several models with process isolation, is free, and is already in your binary. It lacks memory-based eviction (it evicts by count), pins, leases and a refusal with numbers. Compatibility with LM Studio was a design constraint. StudioForge copies its port, /v1/models listing downloaded models, JIT loading, idle TTL, the per-request ttl, the publisher/repo/ folder layout (used in place, so both programs can share one library), the /api/v0/models mirror and the lmstudio://open_from_hf deep link. Moving a client over means changing the host name only. Both programs cannot hold port 1234 at the same time, though, so quit LM Studio or change server.port. What runs where Anything that already talks to an OpenAI-style endpoint needs one base URL and a model id: chat front-ends such as SillyTavern (Chat Completion → Custom), Open WebUI and LibreChat; coding assistants such as Continue, Cline, aider and dsh; the openai SDK in any language; and automation tools that let you type the endpoint. Only the GPU host installs anything. CLIENT MACHINE(S) — NOTHING TO INSTALL Any OpenAI client apps · SDKs · harnesses /v1 MCP agent OpenClaw via sfctl /mcp Browser the control panel :8080 HTTP HTTP HTTP THE GPU HOST — THE ONLY MACHINE THAT INSTALLS StudioForge gateway :1234 one process · registry · VRAM planner · supervisor /v1 OpenAI-compatible API /mcp 19 management tools /api management REST :8080 control panel + tray spawns · proxies, streaming intact Backends — one llama-server per loaded model internal ports 18100–18200 · placed by the planner model A · :18100 one GPU · 32k ctx · 4 slots model B · :18101 two GPUs · 128k ctx restarts the gateway · kills a backend Watchdog sidecar :1235 own process · own MCP server 10 recovery tools answers when the gateway cannot Your GGUF library indexed in place, never copied — models.dir read by each backend PieceWhereWhat it is for Gatewaythe GPU host, one process, 1234/v1, /mcp (19 tools), /api. Holds the registry, the planner and the supervisor. Control panelsame process, second uvicorn, 8080Dashboard, Setup, Models, Download, Chat, Server, Logs. Watchdoga separate process, 1235Its own MCP server with 10 recovery tools. It keeps running when the gateway does not. llama-server children18100–18200, loopback onlyOne per loaded model. A crash takes down one model, never the gateway. Because children bind only 127.0.0.1, the gateway is the only public surface. sfctl companionthe agent's machineA pure HTTP client (Python 3.11+, no CUDA) and the stdio MCP bridge. GGUF librarymodels.dir, where it already isIndexed in place; nothing copied. How a request flows 1 · POST /v1/chat/completions naming a model — or local-model any OpenAI client, on the LAN or a mesh VPN 2 · The registry resolves the id local-model → models.default_model a bad request 404s before any stream starts 3 · Is a backend already serving it? a burst against a cold model makes one load yes Proxy it as-is streaming intact — no planning at all no — it has to be loaded first 4 · The planner picks a placement which GPUs · context, walking the ladder KV cache type · how many parallel slots it may evict idle models to make room or refuse with the numbers — never a CPU spill never a pinned model, a leased card, or one serving 5 · The supervisor spawns it llama-server on a port in 18100–18200 one load at a time; it waits for health the stream is already open: “: loading …”, then “: prefilling …” — no read timeout ever fires 6 · The request is proxied through errors are OpenAI-shaped, with a stable code the idle clock starts when the last request ends 7 · Idles out after its TTL — unless it is pinned a pinned model has no idle timer and is never evicted to make room Everything a client can get wrong is checked before the first byte. A bad model id gets a real 404 with a JSON body, because clients handle that well and routinely mishandle an error frame buried inside a 200 SSE stream. From GET /health on my rig (trimmed): {"status": "ok", "boot": {"phase": "ready", "ready": true}, "engine": {"ok": true, "tag": "b10425", "variant": "cuda", "smoke_tested": true}, "gpu_count": 4, "models_indexed": 34, "can_serve": true} can_serve is the field to poll. status only says the process is alive; can_serve stays false during the first library scan. GET /health?deep=true runs a real 8-token completion against every loaded model, and with nothing loaded it answers no_models_loaded instead of passing. local-model works. local-model, default, auto and current all map to models.default_model, because LM Studio clients fall back to that literal string. A cold load does not look like a hang. The open stream carries : loading <model id> (5s) every five seconds, then : prefilling … until the first token. These are SSE comment lines, which parsers ignore, and they stop client read timeouts from firing. One load at a time, machine-wide. Two cold models planned at once once collided on the same cards and one died out of memory. A request's ttl only moves the idle timer. It can never pin or unpin a model. The VRAM planner This is the part most local servers skip. core/planner.py (3,188 lines in 0.2.0) answers one question before a model starts: given the VRAM free on each card right now, what is the largest window, the best cache quality and the right slot count this model can have? And if none fits, what should the refusal say? The context ladder aim target_ctx: 1,048,576 clamped to trained window halve next rung each rung tries f16, then quantized KV halve floor default_ctx: 8,192 (128,000 on my rig) pass 1 Pass 1: every rung, eviction off any rung fits → load it, evict nothing not even the floor fits Pass 2: every rung again, idle models count as free eviction is paid for → take the highest rung that fits still nothing fits Refused: 507 insufficient_vram, with the numbers never spilled to the CPU an explicit ctx_size is a one-rung ladder, honoured exactly On a first run, tune_for_hardware raises the floor to 16,384 when the smallest card has 24 GiB or more. The planner never offers a window above the model's trained length, because that needs RoPE scaling and costs quality. The second pass exists because of one bad load. There was 79,832 MB free and another 19,423 MB held by one idle model. At that budget, 262,144 tokens on a q4_0 cache would have fitted at 96,004 MB, and 65,536 at f16 at 95,236 MB. What actually loaded was 8,192 tokens at f16, using 89,860 MB. The old code let only the floor evict, so the model paid for the eviction and still got the smallest window. Cache quality is chosen inside each rung: f16/f16 → q8_0/q8_0 → q8_0 K + q4_0 V. Symmetric q4_0 is gone from every automatic path. With a q4_0 K cache, Qwen2.5-7B reproduced only 11.7% of the tokens its f16 self produces, while a q8_0/q8_0 pair measured a KL divergence of 0.0018. KV cache size depends on the architecture WHAT A SLOT'S KV CACHE COSTS, PER LAYER three attention layouts, twelve layers each; schematic, not to scale full Llama, Mistral, most iSWA Gemma 3 / 4 · 1 full in 6 hybrid Qwen3.5–3.8 · 1 KV in 4 full KV cache sliding window (1,024 tokens) recurrent, fixed state The planner reads the layout from the GGUF header; unknown means distrust the numbers Most VRAM calculators compute KV as layers × heads × head-dim × 2 × bytes × context. That is right for Llama and badly wrong for two popular families. Gemma 3 and 4 interleave five sliding-window layers per full layer, at half the head dimension. Qwen3.5, 3.6 and 3.8 declare full_attention_interval = 4, so only every fourth layer has a KV cache and the rest are Gated-DeltaNet layers with a fixed state. Getting it wrong is expensive. Before the fix, the planner estimated 480 GiB of KV alone for a Gemma-4 31B at 262,144 tokens and capped the model at 65,536. For weeks the calibration log had recorded whole-load predictions of 95,615 MB against 40,037 MB actually used. After the fix, the whole load (weights plus KV) is predicted and measured at about 38 GiB at 262,144 tokens across two 5090s. That is a 4× larger window, and every Gemma-4 model benefited. Charging KV to every layer of a Qwen3.5 was the same bug: a flat 4× overcharge. If you write your own estimator, copy llama.cpp's sliding-window cell formula exactly. A flat 1.25× multiplier underestimated by 3.6× at four slots, which means out-of-memory at load. What I measured on four mixed cards The rig is two RTX 5090s and two RTX 3090s, which the panel counts as 111.7 GiB in total, on driver 610.88 (CUDA 13.3). The planner tries one card first, because a split over PCIe without NVLink is slower and runs at the slowest card's pace. Placement (1.5B Q4_K_M, 8k ctx, median of 3)GenerationPrompt processing One 3090352.5 tok/s2,803.6 tok/s Two 3090s, -sm layer344.4 tok/s2,722.5 tok/s Two 3090s, -sm tensor294.3 tok/s1,182.0 tok/s Two 3090s, -sm rowfails: device CUDA2 does not support split buffers One card beat two. Tensor split cost 17% of generation and 58% of prompt processing on this small model, so tensor mode is opt-in and only a benchmark may pick it. -sm row is accepted by the parser and then fails on CUDA, so StudioForge blocks it before spawning. How many parallel slots are worth running is a separate question. Note first that llama.cpp's --ctx-size is the total KV budget shared across slots, not the per-slot window. --ctx-size 4096 with no --parallel reports total_slots: 4, which is 1,024 tokens per conversation. StudioForge launches with ctx_per_slot × parallel. ConcurrentPer streamAggregatep50p95Achieved batch 1302.8 tok/s302.8 tok/s0.41 s0.41 s1.00 2225.3 tok/s425.3 tok/s0.46 s0.49 s1.84 4134.5 tok/s436.0 tok/s0.83 s1.00 s3.46 883.3 tok/s576.9 tok/s1.57 s1.77 s6.03 Qwen2.5-1.5B-Instruct Q4_K_M, one RTX 3090, 8,192 tokens per slot, f16 KV, eight slots launched, 512-token prompts, 192 generated tokens each, 2026-08-19. Reproduced across three runs within 2%. Why is the solo figure here 302.8 when the placement table shows 352.5 for the same model on the same card? They are different benchmark runs with different settings. This sweep launched eight slots of KV and used its own prompt and output lengths. Compare rows within one table, not across the two. The estimator said 8 slots; the measurement said 2. At four slots each stream drops to 44% of its solo speed. The rule picks the largest of 1, 2, 4 or 8 that keeps each stream above 65% of solo while still adding 15% aggregate. Aggregate keeps climbing while each user slows down. Eight slots move 1.9× the tokens of one, but each conversation runs at 27% of solo speed. A rule that maximised aggregate would pick 8 and make every chat three times slower. The batching is real. Achieved batch rising from 1.00 to 6.03 shows shared decode steps, not a queue. Two more flags I measured before trusting them: Speculative decoding helps a single stream only. Qwen3.8-27B Q5_K_S with an MTP head on one 3090, four different prompts: 37.75 tok/s without it, 50.70 tok/s (+34.3%) with draft-mtp at depth 3. Repeating the same prompt showed +751%, which was the prompt cache, not drafting. Above four slots, auto now turns speculation off. A bigger micro-batch trades VRAM for prefill speed. On a 5,166-token prompt, -ub 2048 was 18.6% faster than the default 512 and used 210 MiB more. The planner charges that buffer, rounded up, and raises -ub automatically only above four slots. The refusal, with the numbers Every launch passes --fit off and --n-gpu-layers 999. This matters because llama.cpp build b10425 ships --fit defaulting to on, plus --n-gpu-layers auto (both from PR #16653, December 2025). Together they are a silent partial-offload path, which is a sensible default for a general server and exactly what this project refuses. When nothing fits, you get HTTP 507. A real one, trimmed: HTTP 507 {"error": {"code": "insufficient_vram", "message": "Cannot load '…/gemma-4-31B-it-QAT-Q4_0' entirely in VRAM: needs 29.09 GiB, 20.90 GiB usable. … Suggestions: set KV cache type to q8_0 …; VRAM is held by other processes: …", "studioforge": { "required_bytes": 31235974510, "available_bytes": 22438368871, "per_gpu_free": {"3": 22438368871}, "max_ctx_that_fits": null, "notes": ["wanted up to 262144 tokens of context but not even the 128000 floor fits in the VRAM available right now"], "estimate_mb": {"weights_bytes": 16818.2, "kv_bytes": 8575.0, "compute_bytes": 2438.6, …, "total": 29788.9}, "retry_after_s": null }}} That was the 31B asked to load onto one RTX 3090. The message is for people; error.studioforge is for programs. When a smaller window would fit, max_ctx_that_fits names it. retry_after_s is set only when a busy model is the cause, because otherwise retrying will not help. Pins, TTLs and leases A model loads on first request and unloads after its idle TTL (models.default_ttl_s, swept every 15 s). A pin is a desired state. A pinned model has no idle timer and is never evicted. If it crashes or fails at boot, a reconciler reloads it, backing off from 60 s to 900 s. An explicit unload still wins and keeps it down until someone loads or pins it again. A lease gives a card to one model. A leased card is invisible to every other model's placement, and the owner is forced onto exactly those cards. Asking for a card that is already leased returns 409 conflict, never a takeover. A lease can also hold cards for a program outside the server: reserve_gpus(devices=[3], reason="ComfyUI render") is how I stopped image generation and language models fighting over one card. Eviction never touches a pinned model, one serving a request, one still loading, or a leased card. A background rebalancer may move an idle model to a better placement, but at most once per model per 30 minutes, because a move drops the prompt cache. On my long conversations that cache covered 93% of a 98k-token prompt. When things break VRAM dies with the process that took it. On Windows the children live in an anonymous job object created with JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE, so the kernel kills them however the parent ends. Linux uses PR_SET_PDEATHSIG, which is best effort, so a startup sweep and reclaim_orphan_engines catch what slips through. Why the sweep exists: on 2026-08-18 about 25 GiB across two cards was in use with "everything stopped". The holders were three llama-server.exe children left behind by a test run a coding agent had started and abandoned. Every VRAM holder is now classified as ours, child-of-live-process, orphan, other-instance or foreign, and only orphan is ever killed. Exit codes carry meaning. 2 is a config error naming the key. 3 is a port conflict: the tray waits for the holder instead of respawning into it. 75 means a restart was requested: the server drains and exits (1.0 s measured) and the tray respawns it without counting a crash. That split exists because a GUI restart once left two servers racing for port 1234 and the tray stuck on "Crashed" next to a healthy server. When the gateway is wedged rather than dead, use the watchdog on 1235. It is a separate process built only from argparse and stdlib logging, so it starts even when config.yaml is broken. Its ten tools are health, get_config, set_config, restart_server, kill_model, nuke_all_models, reclaim_orphan_engines, tail_logs, gpu_status and rollback_update. It does not import the code it repairs. The control panel The panel on 8080 runs in the same process and uses no absolute URLs, so it works over plain HTTP or behind an HTTPS proxy. On Windows there is also a tray icon. The dashboard: a 31B split across both 5090s while it processes a prompt, a pinned 26B on one 3090, and every process holding VRAM named with its share. Per-model settings: what the model could do on each set of cards, a fit verdict that re-runs on every change, and every knob to overrule it. Context sizes it cannot reach are greyed out with the reason. The quant picker checks every file against the VRAM free right now before you download it, then refines the badge by reading the GGUF header remotely. The Server tab writes the client configuration for you. Because the MCP entry runs sfctl, the pairing PIN never has to go in a config file. Using it as a harness backend Inference and management are separate paths. Any client uses /v1 with nothing installed. Only an agent that should manage the box needs the MCP tools. Any OpenAI client server.api_key is null by default, so any non-empty key works. Most OpenAI clients refuse to start with an empty one. export OPENAI_BASE_URL=http://my-gpu-rig:1234/v1 export OPENAI_API_KEY=not-required # any non-empty string while server.api_key is unset curl http://my-gpu-rig:1234/v1/chat/completions -H "Content-Type: application/json" \ -d '{"model": "<id from /v1/models>", "messages": [{"role": "user", "content": "hello"}]}' GET /v1/models lists every downloaded model, LM Studio-style, and adds state plus, for loaded models, ctx_per_slot and max_parallel. Ids match as the full publisher/repo/file, a bare filename or publisher/name, case-insensitively. The base-URL setting has a different name in every client, which wastes more evenings than anything else here: Open WebUI uses OPENAI_API_BASE_URL; LibreChat a custom endpoint with baseURL and models.fetch: true; aider OPENAI_API_BASE then --model openai/<id>; Continue provider: openai plus apiBase. dsh takes the same OpenAI provider block as any other server (baseURL, apiKeyEnv); the exact YAML is in the dsh write-up. A keyless server still needs some credential referenced there, because the client insists on a bearer token. OpenClaw on another machine The agent machine installs one small package, sfctl. It needs Python 3.11 and does not depend on the server package, so no CUDA and no planner. sfctl servers add rig http://my-gpu-rig:1234 --api-key <PIN> --use openclaw mcp add studioforge --command sfctl --arg mcp If you edit the config by hand, note that OpenClaw's key is mcp.servers, nested under mcp. The flat mcpServers map is for Claude Code, Cline and LibreChat, and OpenClaw's schema does not know it. Inference goes under models.providers, and the key is baseUrl with a lower-case rl: // ~/.openclaw/openclaw.json { "mcp": { "servers": { "studioforge": { "command": "sfctl", "args": ["mcp"] } } }, "models": { "providers": { "studioforge": { "baseUrl": "http://my-gpu-rig:1234/v1", "apiKey": "not-required", "api": "openai-completions", "models": [ { "id": "<id from /v1/models>", "name": "Rig 27B", "contextWindow": 131072 } ] } } } } The agent sees one merged list of 29 tools in 0.2.0: the gateway's 19 plus the watchdog's 10. Three watchdog tools are renamed recovery_* to avoid clashes; restart_server keeps its name because error messages tell the agent to call it. If the main server is down, the bridge still lists the management tools with a note, so the agent knows they exist. The loop an agent runs: list_models(limit=N): the catalog, newest download first. Read the recommended row. load_model(**row["load_args"]): pass the row through unchanged. load_recommended(model_id, ctx_size=N) when you know the context you need. It refuses instead of shrinking. Run inference over HTTP, not MCP. Naming an unloaded model loads it with planner defaults, not the row you were reading. model_options(model_id) for every context tier, with speeds. search_models → repo_details → download_model to get something new. pin_model for a model that must always answer; reserve_gpus / release_gpus for dedicated cards. server_status and connection_info: what is loaded, who holds VRAM, every address it answers on. Every client also gets this note on connect: INFERENCE IS NOT HERE. This server exposes no chat/completion/generation tool by design. To actually run a prompt, use the OpenAI-compatible HTTP API on the gateway port (POST /v1/chat/completions, /v1/embeddings; GET /v1/models). Naming an unloaded model in a request just-in-time loads it, so you usually do not need load_model at all -- reach for it only to pre-warm a model or to load one with non-default context/quantization settings. Claude Code cannot use this server for its own inference. Its gateway protocols are Anthropic Messages, Bedrock and Vertex, not /v1/chat/completions. It can still use the management tools: claude mcp add studioforge -- sfctl mcp (the -- is required). Benchmark tools should take a lease through POST /api/leases, wait for busy models, and reload what they displaced. My old script ran pkill -f llama-server instead, which kills every backend. What an agent picks from Each catalog row is one real planner call: whether it fits in free VRAM now, on which devices, with how many slots, and whether it would fit if the GPUs were idle (so an agent knows when unloading helps). repo_details reads a GGUF header remotely with HTTP Range requests (2–15 MB, cached) instead of downloading 20 GB, and returns a context-fit matrix from the same planner: Quant1× RTX 50902× RTX 5090All four cards BF16 (51.8 GiB)weights do not fit32k at q8_0256k Q8_0 (27.9 GiB)does not fit256k256k Q5_K_M (19.3 GiB)128k at q8_0256k256k IQ2_M (10.5 GiB)256k256k256k unsloth/Qwen3.8-27B-GGUF, computed on my rig. Figures are the largest window at an f16 cache; "at q8_0" appears only where a quantized cache reaches further. The quant decides the window more than the card count: BF16 to Q5_K_M turns "does not fit" into 128k on one card. To tune a model, benchmark placements, run benchmark_parallel on the winner, then reserve_gpus. Install Windows (the reference platform): Install Git, Python 3.12+, uv and a current NVIDIA driver. Clone the repo or unzip the download. Double-click launchers\Update StudioForge.bat. Despite the name, this is the first-run step: it builds the virtualenv and installs, smoke-tests and pins the newest llama.cpp build your driver supports (you will probably get a newer one than b10425). Run launchers\Start StudioForge.bat, or launchers\StudioForge Tray.bat for the tray icon. The panel opens at http://127.0.0.1:8080 on its Setup tab, where Detect LM Studio library finds your existing models folder. Linux also needs cmake and a CUDA toolkit whose nvcc matches your driver, because upstream publishes no Linux CUDA build and the engine is compiled once per version. As noted above, I have not run this build path end to end myself. git clone https://github.com/LaserLloyd/StudioForge.git && cd StudioForge uv venv --python 3.12 .venv uv pip install --python .venv/bin/python -e ".[dev]" .venv/bin/studioforge serve --open # first run builds the engine For a headless box, deploy/ holds two systemd user units, because the server must run as the user who owns the model library, the venv and the GPU devices. The watchdog unit uses Restart=always and is not bound to the gateway, so it stays up when the gateway is down. Run sudo loginctl enable-linger "$USER" so the units start without a login. ServiceDefault portConfig key Gateway (/v1, /api, /mcp)1234server.port Web control panel8080gui.port Recovery watchdog1235watchdog.port llama-server children (loopback only)18100–18200gateway.child_port_start / _end Config validation refuses a service port that collides with the child range. The data directory is SF_DATA_DIR first, then the folder of a --config file, then <repo>/data; details are in docs/SETUP.md. One instance owns one data directory through an exclusive OS lock. Security The rule: reads, inference and loading stay open; changing the box needs a credential. With server.api_key unset, a request that changes config, restarts, engines, updates, VRAM reclaim, downloads, leases or deletes is accepted only from the same machine, or with the MCP PIN in X-MCP-Pin or as the bearer token. Anything else gets 403 remote_admin_requires_credential. The PIN only guards MCP. It is a pairing code from the startup banner, not an API key. server.api_key is the real credential and ships unset. Set it and it covers /v1, /api, /mcp and the watchdog. All three listeners bind 0.0.0.0 by default. The Setup tab turns its Network exposure row amber while any listener is exposed without a key. Set a key before you expose it. A cross-origin browser request does not count as local, even on loopback. Otherwise any web page could PATCH /api/config at 127.0.0.1:1234. The origin check includes the port and treats Origin: null as foreign. Remote browsers never see the PIN. Image URLs are fetched behind an SSRF guard that blocks loopback, link-local, private, ULA and CGNAT (100.64/10, where mesh-VPN peers live), and connects to the address it vetted. Two limits: a same-machine check trusts anything on loopback, which behind a reverse proxy is the proxy, so put the proxy behind server.api_key. And there is one shared key, with no accounts and no rate limiting. Keep it on loopback behind an authenticating proxy or on a mesh VPN, never on the open internet. Limits Windows is the reference platform. The tray, the job-object VRAM guard and the per-process GPU counters were built there. Linux runs in CI (see Install). macOS is unsupported (no CUDA). NVIDIA only. The planner reads NVML and the engine is a CUDA build. The VRAM figure is an estimate. Weights land within 2% of file size and KV is exact from the per-layer layout, but the compute buffer is a calibrated fraction (clamped to 0.03–0.15, kept in memory, reset by a restart). Multi-GPU splits are proportional, not measured. Interconnect bandwidth is not modelled, and mixed generations run at the slower card's pace. Speed estimates use vendor figures, not measurements. On my rig a dense 31B measured 39.4 tok/s against an estimate of 36.1, and a 122B MoE 37.3 against 47.4. Reserving a GPU for another program only binds my planner. Nothing stops that program, or a third one, from taking the memory. Licence nuance. StudioForge is MIT. pystray, which draws the Windows tray icon, is LGPL-3.0. Installed normally as its own package, that only means keeping its notice. If you freeze everything into one executable, ship the pystray source or keep it a separate module. Nobody else has audited it. Everything above is my description of my own code, so check the source before you trust it with anything that matters. 2,503 unit tests pass and 17 skip in 327 seconds on my machine; CI runs them on Windows and Ubuntu with a fake GPU backend. A second suite loads real models onto real GPUs; it is off by default and needs an environment variable, because of the orphan incident above. Gotchas Quit LM Studio before starting. Both want port 1234. The preflight names the holder; server.port moves StudioForge. Sharing the model folder is fine. --ctx-size is the total across slots, not per slot, and upstream --fit defaults to on. Pass --fit off and size context per slot yourself if you run llama-server directly. Reasoning models can return an empty reply under --reasoning-format auto. Same prompt: content 0 characters and reasoning_content 316 under auto; content 323 under none. reasoning_content is not in the OpenAI schema, so a standard client sees nothing. Use none unless your client reads that field. Do not tree-kill the tray process. Under a venv launcher the watchdog sits in the same process tree, and one restart done that way left the server and the watchdog both dead. Use the tray menu, the panel or sfctl recover --restart. pkill -f llama-server from another tool kills your backends. Take a GPU lease instead. NVML's busId bus number is hexadecimal. "00000000:42:00.0" is bus 66, and reading it as decimal put VRAM on the wrong card without any error. Repeating one prompt in a speculative-decoding benchmark measures the prompt cache. Use distinct prompts. /props also reports speculative.types: "none" while drafting, so read timings.draft_n from a real completion. Vision models get no prompt-cache benefit. llama.cpp disables cache reuse for multimodal models, and each image is budgeted at 1,024 tokens by default, so an 8k window fills fast. Get it GitHub always has the newest version, with the tests, docs, launchers and systemd units: github.com/LaserLloyd/StudioForge git clone https://github.com/LaserLloyd/StudioForge.git The zip is version 0.2.0: main at commit d7e5d26 (2026-08-25), plus prebuilt wheels and sdists for both packages in dist/. Everything in this article (29 tools, the panels, the planner) describes that version. The GitHub main branch has moved on since and adds two management tools, plan_load and check_loaded_model, for 31 in total. SHA-256 of the zip: 72e3b44487645e7790ba8497e029868e07ca340f67c92f1834345d4e50839f6d studioforge-2026-08.zip It holds tracked files only (no config.yaml, no data/). Licence: MIT: use, change, ship or sell it, keep the copyright notice, no warranty. Found a bug? My email is on the About page. The useful parts are the error.code, the numbers from a 507 and the last twenty lines of logs/models/<model>.log. Please leave out the PIN and any key. Pull requests for a joint bin-packing planner, named sampler presets or a working AMD path are very welcome. Where this leaves me Most of this project turned out to be measurement, not code. The planner is arithmetic anyone could write. It became trustworthy only after I read llama.cpp's real KV layout and measured the slot sweet spot at 2 when my estimate said 8. If you already run LM Studio on 1234, trying this costs a clone, one batch file and pointing models.dir at the folder you already have. If it does not earn its place in an afternoon, your old setup is untouched. Related: DeepSeek Harness (dsh) (the harness that calls this server as a plain OpenAI provider), bench-llm / CrucibleForge (the benchmark tool that now takes GPU leases), My OpenClaw Setup (the post that named the problem), and DisPatch (the chat app in front of it).