Skip to main content
Local AI

LLM Quantization 2026: Q2 vs Q4 vs Q5 vs Q8 (K_M, K_S & IQ Explained)

LLM quantization for 2026, explained simply: Q2 vs Q4 vs Q5 vs Q8, the K_M / K_S / IQ variants, and the VRAM-vs-quality trade-off — so you pick the right GGUF for your GPU.

2026 updateThe quantization trade-offs here apply to every current model, including the 2026 lineup (Qwen 3, Llama 4, DeepSeek R1). To see exact VRAM per quant level for any model, use the VRAM Calculator.

Every local LLM tutorial throws around cryptic labels like Q4_K_M, Q5_0, Q8_0. Nobody explains what they mean. Here’s the plain-English version: quantization shrinks model weights from 16-bit floats down to 4 or 2 bits each, trading a little quality for a LOT of VRAM savings.

Quick answer: If you don’t know which quant to pick, choose Q5_K_M — it’s the sweet spot at ~96% of original quality and 40% the file size. Use Q8 if you have plenty of VRAM and want near-reference quality. Use Q4_K_M when VRAM is tight. Avoid Q2 unless you absolutely have to (visible quality drop).

Quantization formats compared (at a glance)

Scanning for the one to download? Here is every common GGUF format side by side. File-size and quality figures are approximate general guidance (they shift a little per model) — not per-model benchmarks.

FormatBitsFile sizeQualitySpeedBest for
Q8_08Largest~99%FastPlenty of VRAM, near-reference output
Q6_K6Large~98%FastMaximum quality on an 8 GB card
Q5_K_M5Medium~96%FastThe default — best balance
Q5_K_S5Medium−~95%FastShave a little VRAM off Q5_K_M
Q4_K_M4Small~94%FastMost popular; tight VRAM
Q4_K_S4Small−~92%FastSqueeze just under a VRAM limit
IQ4_XS4 (i-quant)Smallest 4-bit~93%Slightly slowerSmallest 4-bit at good quality (needs imatrix)
Q3_K_M3Very small~88%FastOnly if Q4 will not fit
Q2_K2Tiny~78%FastEmergency: big model, tiny VRAM
One-line rule: download Q5_K_M unless you have VRAM to spare (go Q6_K/Q8_0) or can’t fit it (drop to Q4_K_M).
Quantization comparison showing Llama 3.1 8B at different bit levels with file size, VRAM requirement, and quality tradeoffs
Same model, seven quantization levels — Q5_K_M is the sweet spot for most people.

The quantization ladder

FP16 — the original, full precision

Size: 14 GB (Llama 3.1 8B) · Quality: 100% · VRAM: 16 GB minimum
What the model was trained as. Only worth running if you have the VRAM and need benchmark-grade quality.

Q8_0 — 8-bit, near-reference

Size: 8.5 GB · Quality: ~99% · VRAM: 10 GB
Effectively lossless. If you have the VRAM, this is the highest reasonable quant for production work.

Q6_K — 6-bit, strong balance

Size: 6.6 GB · Quality: ~98% · VRAM: 8 GB
Barely distinguishable from Q8 in blind testing. Good for 8GB cards that want max quality.

★ Q5_K_M — the sweet spot

Size: 5.7 GB · Quality: ~96% · VRAM: 7 GB
The consensus “best bang for buck” quant. Barely noticeable quality loss, ~40% the size of FP16. If you don’t know which quant to pick, pick this one.

Q4_K_M — 4-bit, most popular

Size: 4.7 GB · Quality: ~94% · VRAM: 6 GB
The de-facto default on HuggingFace. Noticeable but acceptable quality drop. Best when you need to fit a model on tight VRAM.

Q3_K_S — 3-bit, compact

Size: 3.5 GB · Quality: ~88% · VRAM: 5 GB
You’ll notice clunkier responses and occasional coherence issues. Only use if Q4 won’t fit.

Q2_K — 2-bit, emergency mode

Size: 2.8 GB · Quality: ~78% · VRAM: 4 GB
Use ONLY when you absolutely need to fit a big model on small VRAM (e.g., Llama 70B on a 24GB card). Noticeable quality degradation.

What do those cryptic letters mean?

So Q5_K_M = 5-bit quantization using K-quants with medium-tier quality preservation. In 99% of cases you want _K_M variants.

K_S vs K_M: which should you pick?

Short answer: at the same bit level, K_M is the safer default — it keeps the most important weight tensors at higher precision, so it is a touch larger but noticeably steadier in quality. K_S compresses more uniformly: a bit smaller, marginally lower quality. Reach for _S only when _M won’t fit your VRAM.

Q4_K_S vs Q4_K_M: Q4_K_S is a few hundred MB smaller with a small, usually-acceptable quality dip. Pick Q4_K_M unless it spills over your VRAM — then Q4_K_S is the natural step down.

Q5_K_S vs Q5_K_M: both are excellent and the gap is tiny. Q5_K_M is the go-to; choose Q5_K_S only to reclaim a little VRAM without dropping to 4-bit.

IQ vs Q: what are i-quants (and IQ4 vs Q4)?

IQ formats (“i-quants”) use an importance matrix plus non-linear quantization to pack more quality into fewer bits than the older Q..._K scheme. In general they are smaller than the equivalent K-quant at similar quality, but they can be slightly slower on some hardware (older GPUs and CPU-only setups especially) and must be built with an imatrix.

IQ4 vs Q4: IQ4_XS is typically smaller than Q4_K_M at comparable quality — ideal for squeezing a model into VRAM — while Q4_K_M is faster and more universally supported. If it fits, Q4_K_M is the simpler choice; if you are a few hundred MB short, IQ4_XS often saves the day. These are general tendencies, not per-model measurements.

Which quantization should I download?

Your situationRecommended quant
“I have a 24GB card and want maximum quality”Q8_0 or Q6_K
“I want best balance of quality and size”Q5_K_M
“I want to run a bigger model than my VRAM normally allows”Q4_K_M
“I’m fitting a 70B on a 24GB GPU”Q2_K (painful but works)
“I’m fitting a 405B on 80GB”Q4_K_M or Q3_K_M

Common questions

Which quantization should I download?

For most people, Q5_K_M — near-full quality at ~40% the size. Go Q6_K/Q8_0 if you have VRAM to spare, or Q4_K_M if it is tight.

K_S vs K_M — what is the difference?

_M keeps important weights at higher precision (better quality, slightly larger); _S compresses more (smaller, marginally lower quality). K_M is the default; use K_S only to fit VRAM.

Is IQ4 better than Q4?

IQ4_XS is usually smaller than Q4_K_M at similar quality, but Q4_K_M is faster and more widely supported. Pick IQ4 to save VRAM; pick Q4_K_M for speed and compatibility.

Related guides

Last updated: 2026-08-19.

As an Amazon Associate I earn from qualifying purchases. Some links on this site are affiliate links — they cost you nothing extra and never change which product I recommend.