If you’ve spent any time looking for open-source AI models to run in your homelab, you’ve probably landed on Hugging Face. It’s become the de facto place to find everything from small language models to image classifiers. But understanding how Hugging Face actually works—not just how to download from it, but the underlying architecture—changes how you deploy and troubleshoot things. This is what I’ve learned running it for a couple years now.

What Hugging Face Actually Is (And Isn’t)
Hugging Face is not a single monolithic tool. It’s a platform built around a few distinct layers. At the foundation is the Hub, which is essentially GitHub for machine learning models and datasets. Then there’s the Transformers library, which is the Python standard for loading and running those models. And then there are services like Spaces and the Inference API, which let you test models without self-hosting.
When people say “Hugging Face,” they usually mean one of these three things: the Hub website where you browse and download models, the `transformers` library you import in Python, or the services running on their infrastructure. Understanding the difference matters because it affects how you think about self-hosting.
The Hub itself is just a Git-based repository service. Models are stored as Git repos. Datasets are Git repos. It’s straightforward. You clone them, you use them. The Hub provides web hosting, a search interface, and an API for discovery, but the core interaction is Git-based, which is both elegant and surprisingly simple.
The Model Repository Structure and How Data Flows
Open a Hugging Face model repo—say, `mistralai/Mistral-7B-Instruct-v0.1`—and you’ll see the same kinds of files in every repo. There’s a `config.json` that describes the model’s architecture: layer count, hidden dimensions, attention heads, all the hyperparameters. There are the actual model weights, usually in `.safetensors` format these days (safer and faster than pickled PyTorch files). There’s a `tokenizer.json` that defines how text gets converted to tokens. And there’s usually a `README.md` with usage instructions.
When you run `transformers.AutoModel.from_pretrained(“mistralai/Mistral-7B-Instruct-v0.1”)`, here’s what happens under the hood:
- The library makes an HTTP request to the Hub API asking for the repo’s metadata.
- It downloads `config.json` and reads it to determine the model class and architecture.
- It downloads the tokenizer and instantiates it.
- It downloads the model weights—this is the big one, often 5-30GB depending on the model—and loads them into memory.
- It returns a ready-to-use model object.
This is why the first time you load a model takes forever and consumes your bandwidth. Every subsequent load uses a local cache (usually `~/.cache/huggingface/hub/`) so you’re not re-downloading 14GB of weights every time.
The tokenizer bit surprised me early on. I assumed the tokenizer was baked into the model file. It’s not. It’s a separate component that converts your input text into integers that the model understands. Different models use different tokenizers. You can have a fast C++-based tokenizer (handled by the `tokenizers` library) or a slower pure-Python one. For most homelabs, you don’t notice. But if you’re doing batch inference on thousands of texts, it matters.
How the Transformers Library Manages Inference
The `transformers` library is where most of the magic happens for self-hosters. It’s a unified API for loading models of wildly different architectures—LLMs, vision transformers, audio models, multimodal models—and running inference on them. This standardization is worth respecting.
Here’s a basic inference flow with an LLM:
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
model_name = "mistralai/Mistral-7B-Instruct-v0.1"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype=torch.float16,
device_map="auto"
)
prompt = "What is machine learning?"
inputs = tokenizer(prompt, return_tensors="pt")
outputs = model.generate(**inputs, max_length=200)
result = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(result)
The `device_map=”auto”` parameter is important. It tells the library to figure out automatically whether to load the model on GPU, CPU, or split it across both if it doesn’t fit on a single device. This is one of the reasons the library is so useful—it handles the tedious memory management that used to be the homelab bottleneck.
Under the hood, when you call `model.generate()`, the library is handling batching, attention masking, decoding strategy (greedy, beam search, nucleus sampling), and token-by-token autoregressive generation. You’re not writing any of that. That’s why people use `transformers` instead of raw PyTorch.
There’s a catch though: the library assumes you’re either doing inference or fine-tuning. If you want to do something unusual—like extract hidden states from a specific layer, or modify the generation process mid-stream—you end up dropping down to PyTorch directly. This is fine. But it means the abstraction is useful until it isn’t.
Caching, Downloads, and Local Storage
Hugging Face models don’t auto-delete. Once you download them, they live in `~/.cache/huggingface/hub/models—-/`. A 7B parameter model is typically 14GB in float16 precision. Run five different models and you’re at 70GB. This matters in a homelab with limited storage.
You can check your cache with:
huggingface-cli cache-info
And delete old models:
huggingface-cli cache-delete "models--mistralai--Mistral-7B-Instruct-v0.1"
You can also point the cache elsewhere by setting `HF_HOME` environment variable. I keep mine on a separate NAS mount because my main drive fills up constantly.
One thing that got me: if you’re running models in a Docker container, you need to mount the cache directory from the host, or you’ll be re-downloading models every time the container restarts. This sounds obvious in hindsight. I wasted bandwidth learning it the hard way.
For self-hosting, I usually set something like this in my docker-compose:
environment:
- HF_HOME=/models
volumes:
- /mnt/nas/huggingface:/models
This way, multiple containers can share the same model cache, and models persist across restarts.
Spaces and the Inference API: When Not to Self-Host
Hugging Face also offers Spaces, which are like hosted notebooks or containerized apps. You can upload a Gradio or Streamlit app, and they’ll run it for you with generous free compute. There’s also the Inference API, where you send HTTP requests to Hugging Face’s servers and get predictions back.
For homelab purposes, these are mostly useful for testing before you commit to self-hosting. If you want to know whether a model is useful for your use case, spin up a Space first. Run it for an hour. See if it’s actually what you need. Then, if it makes sense, download the model weights and run it locally where you have full control.
The Inference API pricing is reasonable for one-off experiments but gets expensive fast if you’re doing serious inference work. That’s when you self-host. The Spaces tier limits are also something to watch—they’ll put you in a queue if you exceed free usage. Not ideal for production.
Dependencies and the Python Ecosystem Problem
Here’s where things get a bit frustrating. The `transformers` library depends on PyTorch (or TensorFlow, but PyTorch is dominant), which has its own GPU acceleration layers. The quality of GPU support varies wildly by hardware. NVIDIA GPUs have great support via CUDA. AMD GPUs work but are more fragile (ROCm is improving but it’s not there yet). Intel Arc is getting support but it’s recent. And if you’re on an older NVIDIA card, you might be stuck on an older CUDA version.
I spent three days last year trying to get GPU acceleration working on an RTX 3080 after a Ubuntu update broke my CUDA setup. The solution was downgrading and pinning PyTorch version. This shouldn’t be a surprise anymore, but software dependency hell is still real.
For CPU-only inference, `transformers` works fine but is slow. A 7B model on modern CPU hardware takes 2-5 tokens per second. It’s functional but not pleasant.
There’s also the quantization layer. Hugging Face doesn’t enforce quantization, but the ecosystem does heavily use it. You can load a model in 4-bit quantization using `bitsandbytes`, which reduces the effective memory footprint from 14GB to 4GB. This works well. But it’s an additional library, additional configuration, and occasionally a source of bugs.
Authentication and Private Models
If you upload a private model to Hugging Face, you can still download and use it via the library. You just need to authenticate first:
huggingface-cli login
This stores a token in `~/.huggingface/token`. If you’re running in Docker or in a CI/CD pipeline, you pass the token as an environment variable:
HF_TOKEN=hf_xxxxx python your_script.py
This is straightforward enough, but worth noting: the token has broad permissions. If you’re sharing a homelab machine with others, be careful where you store it. And if you’re running inference in a containerized setup, the token lives in the container, which has its own security implications if you’re not careful with image layers.
The Broader Picture: Why This Matters
Understanding how Hugging Face works isn’t just academic. It explains why your first inference is slow (network and downloads). It explains why you need to manage cache carefully (models are large). It explains why GPU support matters (CPU inference is slow). And it explains the value proposition: a unified API across completely different model types, with a living ecosystem of pre-trained models that people maintain and improve.
The platform works because it solves a real problem that existed before: you used to have to know PyTorch or TensorFlow deeply to run any model. You had to understand tensor shapes, device management, optimization tricks. Now you don’t. You just download and run. That’s the core win, and it’s worth the abstraction layers it introduces.
If you’re building a homelab setup with local AI inference, Hugging Face is doing the heavy lifting. You’re not managing model architectures or writing optimization code. You’re just writing the application logic that sits on top. That’s worth understanding, even if you never look at the source code.
FAQ
How much storage does Hugging Face need?
A single 7B parameter LLM is typically 14GB in float16 format. Quantized versions drop to 4-7GB. Cache can grow to 100GB+ if you work with multiple large models. Budget accordingly on your homelab storage.
Can I run Hugging Face models without GPU?
Yes, but CPU inference is slow—expect 2-5 tokens per second on modern CPUs for a 7B model. GPU acceleration is strongly recommended for anything production-like in your homelab.
Do I need to download entire models?
Yes, the library downloads the full model weights. You can’t partially load models. If you need to save bandwidth, use quantized versions (4-bit or 8-bit) or smaller model variants designed for mobile/edge.
Is Hugging Face self-hostable?
The Hub itself is not self-hostable, but the actual models are. Download them once, cache them locally, and they run completely offline. Only the initial discovery and download require internet.
What’s the difference between float16 and quantized models?
Float16 models use 16-bit precision and are larger but higher quality. Quantized models (4-bit, 8-bit) use lower precision, take less memory, run slightly faster, with minimal quality loss for most use cases.
Explore Hugging Face in our AI Homelab Toolkit.
