I spent three months paying for Claude API credits before I realized I could just run models locally. The wake-up call came when I burned through $40 in a weekend testing a script I’d written. That’s when I found Ollama, and honestly, the install process is the only reason it took me a day to get running instead of an hour—not because Ollama itself is complicated, but because I was overthinking it.

This walkthrough covers installing Ollama on Linux, running your first model, and avoiding the small mistakes I made along the way.
The Problem and Why Ollama Solves It
Running language models locally used to mean wrestling with Python environments, CUDA versions, and half a dozen dependencies that broke every time you updated your system. You’d spend two hours getting Llama 2 running in llama.cpp, then realize you needed a different quantization or wanted to try Mistral, and you’d start over. Or you just paid OpenAI, which works but costs money and sends your data to their servers.
Ollama wraps all of that friction into a single binary. Download it, run one command, and you have a working LLM with an API server listening on localhost:11434. It handles model downloading, VRAM management, GPU acceleration if you have it, and even quantization. You don’t have to think about any of it.
Prerequisites and Hardware
You’ll need a Linux system. This guide covers Ubuntu/Debian; if you’re on Fedora or Arch, the package manager steps change but nothing else does. I’m assuming you have root or sudo access.
For hardware, here’s what matters:
- Minimum RAM: 8 GB for small models (Phi 2.7B, Mistral 7B). Larger models (Llama 70B) will need 64+ GB or significant swap, which is slow.
- GPU (optional but faster): NVIDIA with CUDA support is easiest. AMD cards work via ROCm. Intel Arc works. CPU-only is viable for smaller models but expect a 5-10x speed hit.
- Disk space: Models range from 3 GB (Phi) to 40+ GB (Llama 70B). Have at least 100 GB free if you plan to experiment with several.
- CPU: Any modern processor works. Doesn’t need to be a powerhouse if you have a GPU.
I’m running this on an Ubuntu 22.04 box with 32 GB RAM and an RTX 4060 Ti. Models load in seconds, inference runs at about 40-50 tokens/sec for 7B models. Your mileage varies.
Step 1: Download and Install Ollama
Head to https://ollama.com and grab the Linux installer, or use the curl command:
curl -fsSL https://ollama.ai/install.sh | sh
This downloads the binary, sets up a systemd service, and creates an ollama user. It takes about a minute. When it finishes, Ollama is already running as a service in the background.
Check that it’s working:
systemctl status ollama
You should see active (running). If not, check the logs with journalctl -u ollama -n 50 to see what went wrong. Usually it’s a permission issue or a missing dependency like libc.
The binary lives at /usr/bin/ollama and the service runs as the ollama user. Models are stored in /usr/share/ollama/.ollama/models by default. If you want to move that location (say, to a larger disk), you can set the environment variable OLLAMA_MODELS before starting the service, but for now leave it alone.
Step 2: Install GPU Support (Optional)
If you have an NVIDIA card, the installer usually detects and installs CUDA drivers automatically. Verify with:
nvidia-smi
If that works and shows your card, Ollama will use it. Restart the service to be sure:
sudo systemctl restart ollama
For AMD cards with ROCm, the setup is more involved. You’ll need the ROCm runtime and appropriate libraries. Follow the official AMD docs, then set HSA_OVERRIDE_GFX_VERSION if needed. It works, but it’s less polished than NVIDIA support.
If you have neither, don’t worry. CPU inference is slow but it works. A 7B model on a modern CPU will generate output, just not quickly. Good for experimenting; not good for production use.
Step 3: First-Run Configuration and Pulling Your First Model
Now that Ollama is running, you pull a model. This downloads it and makes it available. The easiest starting point is Mistral 7B, which is fast and coherent:
ollama pull mistral
This takes a few minutes depending on your internet speed. The model is about 4.1 GB compressed, and Ollama decompresses it on disk. You’ll see a progress bar. When it finishes, you have a working LLM.
Try running it:
ollama run mistral
This starts an interactive chat session. Type a prompt and hit Enter:
>> Tell me about rust programming in 2 sentences
Rust is a systems programming language that emphasizes safety,
concurrency, and performance by preventing common mistakes like null
pointer dereferences and buffer overflows at compile time. It has grown
in popularity for building reliable, efficient software in domains like
web servers, operating systems, and embedded systems.
Type exit or press Ctrl+D to quit the chat.
By default, Ollama listens on 127.0.0.1:11434. If you want external access, set OLLAMA_HOST=0.0.0.0:11434 in the systemd environment. Edit /etc/systemd/system/ollama.service (or create an override with systemctl edit ollama) and add:
[Service]
Environment="OLLAMA_HOST=0.0.0.0:11434"
Then reload and restart:
sudo systemctl daemon-reload
sudo systemctl restart ollama
Step 4: Using the OpenAI-Compatible API
The real power of Ollama is that it runs an API server compatible with OpenAI’s format. Any tool expecting ChatGPT can point at Ollama instead.
Test it with curl:
curl http://localhost:11434/api/generate -d '{
"model": "mistral",
"prompt": "Why is Rust popular?",
"stream": false
}'
You get back JSON with the response. If you use the chat completions endpoint:
curl http://localhost:11434/v1/chat/completions
-H "Content-Type: application/json"
-d '{
"model": "mistral",
"messages": [{"role": "user", "content": "Hello"}]
}'
This works with any OpenAI-compatible client. If you’re using Python with the OpenAI library, just point it at Ollama:
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:11434/v1",
api_key="ollama" # dummy key, ollama doesn't check
)
response = client.chat.completions.create(
model="mistral",
messages=[{"role": "user", "content": "Explain Docker in one sentence"}]
)
print(response.choices[0].message.content)
Same for Node.js, Go, or anything else. This is why Ollama is useful—you write code once and it works whether you’re using OpenAI, Ollama, or LM Studio.
Common Issues and Fixes
“Connection refused” when trying to access the API: Make sure the service is running (systemctl status ollama). Check that you’re not blocked by a firewall. If Ollama is on a different machine, verify OLLAMA_HOST is set to 0.0.0.0 not 127.0.0.1.
“Out of memory” errors: Ollama will try to load the entire model into VRAM if possible, otherwise it uses system RAM. If you’re running a large model on a small GPU, it degrades gracefully to CPU, but slowly. Try a smaller model first. Phi 2.7B uses about 5 GB of VRAM; Mistral 7B uses 15-20 GB depending on quantization.
Models aren’t using GPU: Run ollama list to see what’s loaded, then check if GPU is available with nvidia-smi. Restart the service if you just installed drivers. Check the logs with journalctl -u ollama for more detail.
Slow generation: If you’re getting 1-2 tokens/sec, you’re on CPU. Pull a smaller model or install GPU support. Some quantization levels are slower—try the default (usually Q4_K_M) before fine-tuning.
Disk space fills up quickly: Models are decompressed on disk. A 7B model in Q4 format takes about 4 GB. If you’re pulling many models, specify a different OLLAMA_MODELS directory pointing to a larger drive before pulling.
What to Do Next
Once you have Ollama running, try pulling other models. Llama 2 7B is solid for general tasks. For coding, CodeLlama or Mistral work well. For pure speed on limited hardware, Phi 2.7B is surprisingly capable.
If you’re building something more permanent, containerize it. Ollama works fine in Docker; just mount the models directory to persist them across container restarts. Or integrate it with other tools—n8n can call Ollama as a step in a workflow, or build a simple web UI in front of the API.
The one thing that surprised me after using this for a few months is how much you start forgetting about the model itself. You’re not choosing between Claude and GPT anymore—you’re just picking whatever quantization of whatever model fits your hardware, and it becomes background infrastructure, like running a database. That’s when you know a tool is doing its job right.
FAQ
Can Ollama run on a Raspberry Pi?
Yes, but slowly. Ollama supports ARM64, so it installs on newer Raspberry Pi 4 and Pi 5. For anything useful, you’ll want 8 GB of RAM and patience—even small models like Phi run at 1-2 tokens/sec. A Pi 5 with 8 GB is workable for experiments; don’t expect production speed.
How much RAM do I need for Ollama?
For small models (Phi, Mistral 7B), 8 GB minimum. The model loads into memory, so a 7B model in Q4 format needs roughly 4 GB VRAM or system RAM plus overhead. Larger models (13B, 70B) need 16+ GB. GPU VRAM helps but isn’t required.
Does Ollama send data to the cloud?
No. Everything runs locally on your machine. No telemetry, no cloud calls, no API keys. The API listens on localhost by default, though you can expose it to your network if needed. Your prompts and responses never leave your hardware.
What’s the difference between Ollama and LM Studio?
Ollama is a headless service with a CLI and API; LM Studio has a GUI. Both run models locally. Ollama is lighter and better for servers or automation; LM Studio is better if you want a visual interface. They’re not mutually exclusive—use whichever fits your workflow.
Can I use Ollama with Docker Compose?
Yes. You can run Ollama in a container, though it’s often simpler to run it natively on the host. If you containerize it, make sure to mount the models directory to a persistent volume so models survive container restarts, and pass through GPU support with runtime: nvidia in the compose file.
Explore Ollama in our AI Homelab Toolkit.
