I do not keep every local-AI experiment running. Most of them were useful for a weekend and then became another container I had to patch. What stays on is a small local LLM homelab stack: one inference path, a short model list, a chat UI I actually open, and a hardware path I can explain without a spreadsheet.
This is the Builders inventory β not a lab tour of everything I tried.
The stack that stays powered on
Inference: Ollama in Docker
Ollama is still the front door. I run it in Docker with GPU passthrough, persistent model volume, and the OpenAI-compatible API on port 11434. Other tools in the lab talk to that endpoint. I am not juggling three competing runtimes for day-to-day chat and scripts.
I have used llama.cpp and LM Studio for checks. They are fine. They are not what I leave running overnight. Ollama is the one that stays because the rest of the lab already expects it.
Chat UI: Open WebUI beside it
Raw CLI is enough for a smoke test. Daily use needs a browser UI. Open WebUI sits next to Ollama, talks to the same API, and gives me model switching without teaching every visitor a curl one-liner. When I want character-heavy sessions I still reach for SillyTavern, but that is optional β not part of the always-on core.
Models: a short list, not a museum
I keep a small set that matches the VRAM I actually have:
- A fast 7β8B class model for everyday chat and quick scripts
- A stronger mid-size model when quality matters more than speed
- One coding-oriented model when I am rubber-ducking Compose or Ansible
I delete models I have not loaded in weeks. Disk is cheap until the volume is full of quant variants you will never open again. For which sizes fit which cards, I send people to the Local LLMs hub instead of repeating matrices in every post.
Hardware path: VRAM first, then the box around it
The stack only makes sense if the card can hold the model. My shopping order is still:
- Pick the model class you care about (8B vs 14B vs 32B+).
- Match VRAM (and system RAM for CPU offload cases).
- Then worry about case, PSU, and noise.
That is why I keep the free local LLM hardware checker on the site β GPU + RAM in, fit list out β and why Homelab Gear lists the cards and mini PCs I actually point people at. I am not publishing fake tokens/sec race tables here. Real speed depends on quant, context length, and backend.
What I turned off on purpose
- Extra inference servers βjust in caseβ
- Every new UI that duplicates Open WebUI
- Cloud proxy layers I do not need for private lab chat
- Oversized models that thrash VRAM so hard the box is unpleasant to use
The rule is boring: if it is not in the weekly path, it does not get a restart policy.
How the pieces talk to each other
Request flow on a normal day:
- Browser β Open WebUI (or a script).
- UI / script β Ollama API (localhost:11434 on the lab network, reverse-proxied when I need it from another machine).
- Ollama β GPU for the loaded model.
- Optional: Home Assistant or small Python jobs hit the same API for summaries and automations.
One API surface. Multiple clients. That is the whole architecture pitch.
Where Homelab Gear and the hubs fit
If you are choosing the next card or a quiet mini PC, start on Homelab Gear. If you need the model and setup map, use the Local LLMs hub. If you only want a fast answer for your GPU and RAM, open the hardware checker first.
Those three pages are the soft CTA for this stack. No signup wall on the checker. No claim that my tokens/sec will match yours.
Soft next step
Keep the stack small. Ship Ollama + one UI + models that fit. Validate the card with the checker, then buy from the Homelab Gear path only when the fit list says you are stuck. That is the local LLM homelab stack I still leave on.
FAQ
What is the minimum I should run?
Ollama with GPU passthrough, one 7β8B model that fits your VRAM, and a simple chat UI. Everything else can wait.
Do I need a 24GB card on day one?
No. An RTX 3060 12GB-class card is enough for the everyday 7β8B path. Step up when you know you need 14B+ quality or larger context.
Where should I check my current PC first?
Use the free local LLM hardware checker, then the Local LLMs hub and Homelab Gear if you are shopping.
