Skip to main content

LM Studio Not Starting? Common Errors and Fixes

LM Studio crashes on launch or won't load models? Walk through the most common startup errors, GPU issues, and config problems I've hit running local LLMs.

LM Studio crashes on startup. Or it opens, you download a model, and the whole thing freezes. Or worse: it loads fine but refuses to actually run inference, sitting there with a spinning wheel while your GPU sits idle. I’ve been there. Over the last few months running LM Studio locally alongside other homelab services, I’ve hit enough of these walls to recognize the patterns. Most of them have fixes that take five minutes once you know what you’re looking for.

๐ŸŽฏ Not sure if this will run on your hardware?Use our free Local LLM Hardware Checker โ€” pick your GPU and RAM, see which models will run with real tokens/sec estimates.
Check my hardware โ†’
LM Studio screenshot
LM Studio u2014 from the official site

CUDA Out-of-Memory on Load

This one hit me first. Downloaded a 7B parameter model, clicked load, and got a vague “CUDA out of memory” error. The frustrating part: I have a 12GB GPU. The model should fit.

The issue was quantization. LM Studio defaults to loading GGUF models at certain precision levels. A Q5_K_M quantized 7B model takes around 5GB of VRAM, but the unquantized version or a higher-precision variant can push 11-12GB. I was downloading the wrong file.

Check your model’s Hugging Face card before downloading. Look for the quantization suffix: Q4_K_M and Q5_K_M are the safe bets for mid-range consumer GPUs. Q6_K_M and Q8_0 need more headroom. If you’re running on less than 8GB VRAM, stay with Q4 or lower.

Once I switched to a properly quantized variant, it loaded in seconds. LM Studio’s model browser doesn’t always make the quantization obvious at first glance, so you have to click through to the actual Hugging Face repo to see what you’re getting. That’s the annoying part. The fix itself is just pick a smaller file.

Advertisement

GPU Not Detected on Windows

Windows was a surprise. Downloaded LM Studio on a 3070 Ti machine, launched it, and the app said “No GPU detected, running on CPU.” CPU inference on a 13B model is unusable. Maybe 15 seconds per response.

The culprit: NVIDIA CUDA drivers were outdated. LM Studio needs CUDA 11.8 or higher (depending on the version). I had 11.4 sitting there. Windows Update hadn’t touched GPU drivers in months.

Went to the NVIDIA driver download page, grabbed the latest, rebooted. LM Studio saw the GPU on the next launch. Check your driver version by opening NVIDIA Control Panel or running nvidia-smi in cmd if you have the CUDA toolkit installed. If you don’t have the toolkit, grab the GPU drivers directly from NVIDIA’s site, not Windows Update.

One caveat: AMD GPUs have a different path. If you’re on AMD, you need to download models that support ROCm, and LM Studio support for AMD is newer and less stable than NVIDIA. I haven’t had good luck with it personally, but it’s worth trying if that’s your setup.

Server API Port Already in Use

LM Studio runs a local API server on port 1234 by default. I started using it to hook models into Home Assistant automation, and suddenly LM Studio wouldn’t start the server. Kept getting “port 1234 already in use.”

First instinct: something else was running on that port. Checked with netstat -ano | findstr 1234 on Windows (or lsof -i :1234 on Mac/Linux). Nothing. Rebooted the machine. Still the same error.

Turned out to be a Windows socket state that didn’t clear on reboot. The port was technically available but Windows still thought it was bound. Waited 15 minutes and tried again. It worked. If you hit this, try changing the port in LM Studio’s settings first. Go to Server > Change Port and pick something like 1235. Takes two seconds and you can move on without waiting for the OS to forget.

You can also configure it in the advanced settings to use a different port on startup, but the UI is easier if you’re in a hurry.

Model Loading Hangs Indefinitely

Download a model, click load, and the progress bar gets stuck at 87%. Stays there. Doesn’t crash, doesn’t error, just hangs. I left one running for 45 minutes thinking it was slow. It wasn’t coming back.

This usually means a disk I/O bottleneck or a corrupted download. LM Studio caches models in a local directory (usually ~/.cache/lm-studio on Linux/Mac or C:Users[username]AppDataLocalLM Studio on Windows). If the model file is incomplete or corrupted, loading fails silently.

Delete the partial model from the cache directory and re-download. LM Studio should pick it back up. If the download keeps failing partway through, check your disk space first. A full drive will kill the download without a clear error message. Also check your internet connection. Some models are 4-7GB, and a unstable connection will corrupt the file.

If it’s a persistent issue with one model, try a smaller variant or a different model from the same creator. Sometimes the model file itself is bad on Hugging Face.

Chat UI Responds But Inference Never Completes

This one is subtle. LM Studio starts, the model loads, you type a prompt, and it sits there. The chat window shows a cursor blinking. No error. No crash. Just infinite waiting. Sometimes it takes 20 seconds, sometimes it genuinely hangs.

Two common causes: context window too large or GPU memory getting thrashed by other processes. LM Studio’s default context window on some models is set high. A 4096-token context on a 7B model can use most of your VRAM, leaving nothing for the actual inference computation. Sounds backward, but it’s how it works.

Try reducing the context length in LM Studio’s settings. Start with 2048 or even 1024 tokens. Run a quick inference. If it completes in a reasonable time, you’ve found your ceiling. You can bump it up incrementally from there.

The other thing: close everything else. Chrome with 30 tabs running, VS Code, Slack, whatever. GPU memory fragments, and CUDA gets confused about where to put things. Sounds paranoid, but I’ve watched inference go from hanging to 4 seconds per response just by closing unnecessary apps. Especially if you’re running on a mid-range GPU under 12GB total.

Settings Revert After Restart

Changed the context window, set the API port, selected my GPU. Closed LM Studio. Opened it again. Everything back to defaults. Happened twice before I figured it out.

LM Studio stores settings in a config file, but if the app doesn’t have write permissions to that directory, changes don’t persist. This mostly happens on fresh installs or if the config directory got read-only permissions somehow.

On Windows, right-click the LM Studio folder, check Properties, and make sure it’s not marked Read-Only. On Linux/Mac, check permissions on ~/.cache/lm-studio and ~/.config/lm-studio (if that exists). Should be owned by your user with 755 or 775 permissions. One chmod command fixes it usually: chmod -R u+w ~/.cache/lm-studio.

After that, settings stuck. Small thing, but annoying until you know where to look.

What Actually Matters When Troubleshooting

If LM Studio isn’t working, most of the time it’s one of these five things. Start from the top: VRAM pressure, driver version, port conflicts, incomplete downloads, or context size. Most homelab setups work fine once you’ve hit all of these once and learned the boundaries. The app itself is pretty solid. It’s the boundary conditions that get you.

One thing I wish I’d known earlier: keep a note of what model you’re actually downloading. The Hugging Face filenames are dense and easy to mix up. I grabbed a GGUF that was designed for another framework once and wasted 20 minutes wondering why loading was slow. Read the model card. One minute of reading saves 20 minutes of troubleshooting.

Advertisement

FAQ

How much VRAM do I need to run LM Studio?

At minimum, 6GB for small 7B models in Q4 quantization. 8GB gets you flexibility. 12GB or more lets you run 13B models comfortably. CPU-only mode works but is slow enough that you probably won’t use it.

Can LM Studio run on Mac with Apple Silicon?

Yes. Download the Mac version from the website. It uses Metal acceleration instead of CUDA, and performance is actually respectable. Much better than the first versions.

Does LM Studio work with AMD GPUs?

It supports ROCm for AMD hardware, but the integration is newer and less tested than NVIDIA support. If you’re on AMD, expect more friction and fewer model options.

Can I use LM Studio models with other applications?

The API server exposes a standard interface compatible with OpenAI-compatible clients. Any app that can point to a custom API endpoint can use LM Studio models.

What’s the difference between GGUF models and other formats?

GGUF is the format LM Studio uses natively. It’s quantized and optimized for local inference. Other formats like SafeTensors need conversion. Stick with GGUF for LM Studio.

Explore LM Studio in our AI Homelab Toolkit.

Written by Engineer running a 24/7 homelab since 2022. Every guide here is built and tested on my own hardware. No paid placements.

Keep reading