Quantization Degrades LLM Knowledge Non-Linearly: 2026 Guide

quantizationggufllmselfhosted

TL;DR: A 55-quantization study of Qwen3.6 27B (Quesma, August 2026) found factual knowledge survives quantization down to 5-bit almost untouched, drops noticeably at 3-bit, and falls off a cliff at 2-bit — a non-linear pattern, not a smooth slide. Q4_K_M sits just above the cliff and remains the sane default. Below Q4, you are trading away stored facts faster than you are saving VRAM.

Q5_K_M and upQ4_K_MQ3 / Q2
Best forKnowledge-heavy work without RAGDefault for chat, coding, agentsAlmost nothing at 27B scale
VRAM for a 27B dense model~20 GB+ file, needs 24 GB+ card~17 GB file, fits 24 GB with headroomFits 16 GB, at a real accuracy cost
The catchMarginal gain over Q4 on most tasksMeasurable loss on obscure factsSteep, disproportionate knowledge loss

Honest take: run the biggest model you can hold at Q4_K_M or better; if a model only fits your card at Q3 or below, pick a smaller model at Q4 instead — the quantization cliff costs you more knowledge than the parameter count gives back.

What did the Qwen3.6 27B quantization study actually find?

Factual knowledge degrades non-linearly with bit-width: quantizations of Qwen3.6 27B at 5-bit or above (file sizes over 20 GB) scored the same as the original BF16 weights on a factual-recall benchmark, 3-bit variants dropped noticeably, and 2-bit variants dropped steeply. That is the core result of the Quesma study (“Quantization hurts knowledge nonlinearly,” August 2026) that circulated on r/LocalLLaMA.

The methodology is worth spelling out, because it is unusually thorough for a community benchmark:

  • 55 GGUF quantizations of the same model, Qwen3.6 27B, pulled from the three main sources self-hosters actually use: Unsloth, Bartowski, and llama.cpp’s own reference quants.
  • Benchmark: Incompressible Knowledge Probes (IKP), a tiered trivia benchmark built so the answers cannot be reasoned out — the model either stored the fact or it didn’t. Tiers run from common knowledge (T1) to obscure facts (T7); the unquantized model itself spans roughly 99.5% accuracy on T1 down to 4% on T7, so the benchmark has headroom in both directions.
  • Finding 1 — a threshold, not a slope: every quant above the ~20 GB / 5-bit line matched BF16 on factual recall. The money you spend on Q6_K or Q8_0 buys you nothing on this benchmark over Q5_K_M.
  • Finding 2 — obscure facts go first: the drop at 3-bit and 2-bit concentrates in the harder tiers. A Q2 model still answers “what is the capital of France” and still sounds fluent; what it lost is the long tail — exactly the knowledge you cannot easily spot-check in casual use.
  • Finding 3 — KL divergence tracks the damage: the degradation correlates linearly with the quant’s KL divergence from the BF16 model’s output distribution. Bit-width is a poor predictor; divergence from the original model is a good one.

That last point matters practically. Two quants with the same nominal bit count (say, a Q3_K_M from two different makers, with different imatrix calibration) can land at different KL divergence — and the one closer to the original model keeps more of its knowledge. “Q4” is not one thing; it is a family of trade-offs.

Quesma ran a follow-up on Qwen3.8 27B in September 2026 (“4-bit holds up, 1-bit collapses”) and saw the same shape on the newer model: 4-bit intact, collapse at the bottom of the bit range. Two models, same cliff.

Why is the degradation non-linear?

Because knowledge storage in an LLM is dense and low-redundancy, and quantization noise crosses a threshold where rare facts become unrecoverable. The same study points out the contrast: across model sizes, knowledge scales close to linearly — a 14B model knows roughly proportionally less trivia than a 27B. Across bit-widths, it doesn’t. Accuracy is flat from 16-bit down to 5-bit, then bends, then dives.

The intuition most consistent with the data: a frequently-reinforced fact (Paris, France) is encoded redundantly across many weights, so it survives a lot of rounding noise. A fact the model saw a handful of times in training lives in a few precise weight values. Round those aggressively and the fact doesn’t get fuzzy — it disappears. That is why the loss concentrates in IKP’s upper tiers rather than spreading evenly, and why perplexity (which is dominated by common patterns) under-reports the damage. A Q2 model’s perplexity looks merely “worse”; its long-tail recall is gutted.

You can measure where a specific quant sits yourself. llama.cpp ships KL-divergence tooling in llama-perplexity:

# 1. Record reference logits from the highest-precision GGUF you have
./llama-perplexity -m qwen-27b-bf16.gguf -f wiki.test.raw \
  --kl-divergence-base qwen-27b-logits.bin

# 2. Score any quant against that reference
./llama-perplexity -m qwen-27b-Q3_K_M.gguf \
  --kl-divergence-base qwen-27b-logits.bin --kl-divergence
# → reports mean KLD; per the study, lower KLD ≈ more knowledge retained

Given finding 3 above, mean KLD against the full-precision model is the best single number you can get locally for “how much did this quant cost me” — far more informative than the bit label in the filename.

Which quantization level is safe for each VRAM tier?

For a 27B-class dense model, Q4_K_M (~17 GB file) is the floor worth running, and it needs a 24 GB card to run well. Below that line, drop the parameter count, not the bits. Approximate GGUF file sizes for a 27B dense model, with the study’s results mapped onto common cards:

Your cardFits (27B dense)Knowledge verdict (per IKP results)Better move
48 GB+ (or 2×24 GB)Q8_0, ~29 GBIndistinguishable from BF16Spend the surplus on context, not bits
32 GB (RTX 5090)Q6_K, ~22 GBIndistinguishable from BF16Q6_K + long context is the sweet spot
24 GB (RTX 3090 / 4090)Q4_K_M, ~17 GBJust above the cliff; minor long-tail lossQ5_K_M (~19–20 GB) if you run short contexts
16 GBQ3 onlyNoticeable factual loss — below the safe lineRun a 12–14B model at Q4/Q5 instead
12 GBQ2 onlySteep loss; worst of both worldsRun a 7–8B model at Q5/Q6 instead

Two things this table bakes in:

Bigger-at-Q4 beats smaller-at-Q8 — down to the cliff, and not past it. Since knowledge scales near-linearly with parameters but holds flat from 16-bit to 5-bit, a 27B at Q4_K_M knows more than a 14B at Q8_0 in the same ~17 GB footprint. The old rule of thumb survives — but only above the cliff. A 27B at Q2 to “fit” a 12 GB card inverts it: you keep the parameter count on paper and lose the knowledge it was supposed to carry.

File size is not VRAM usage. The KV cache sits on top of the weights and grows with context length. A 17 GB Q4_K_M on a 24 GB card leaves ~6 GB for cache and buffers — fine at 8K context, tight at 32K. Our context window guide covers the math per model.

If you’re deciding between a 24 GB and a 32 GB card over exactly this question, rent both for an evening before spending four figures — an RTX 3090 runs from $0.07/hr and a 4090 from $0.14/hr on Vast.ai (marketplace pricing, September 2026), which is enough to load your actual model at Q4 and Q6 and diff the answers on your own prompts. Used 3090s ran $1,150–$1,350 in September 2026; that evening of testing is cheap insurance.

A problem you will actually hit: the OOM that pushes you below the cliff

The most common way self-hosters end up at Q3 is not a deliberate choice. It goes like this: Q4_K_M runs fine in short chats, then you raise the context to 32K for a long document, the KV cache blows past your remaining VRAM, the runtime OOMs or crawls under CPU offload — and the “fix” everyone reaches for is downloading the Q3_K_M, because it’s the next size down. That fix quietly moves you onto the steep part of the knowledge curve.

The better fix keeps the weights at Q4 and shrinks the cache instead. In llama.cpp (and anything built on it, Ollama included), quantize the KV cache rather than the weights:

./llama-server -m qwen-27b-Q4_K_M.gguf -c 32768 \
  -ctk q8_0 -ctv q8_0 -fa on

q8_0 KV roughly halves cache memory versus the f16 default with little measurable quality cost — a far better trade than dropping weight precision from 4-bit to 3-bit, because the study shows the weights are where the knowledge lives. Flash attention (-fa on) is required for quantized V cache in llama.cpp as of late 2026 builds.

Does the cliff appear in GPTQ and AWQ too, or is it GGUF-specific?

The cliff is not a GGUF artifact — it shows up across quantization methods, though the methods differ in where exactly they fall off. “An Empirical Study of Qwen3 Quantization” (arXiv 2505.02214, May 2025) tested five post-training quantization methods — RTN, GPTQ, AWQ, SmoothQuant, and BiLLM — on Qwen3 models from 0.6B to 72B at bit-widths from 1 to 8. The pattern rhymes with the GGUF results:

  • AWQ holds at 4-bit, degrades rapidly below it. At 4-bit AWQ, drops versus 8-bit were modest across all model sizes; at 3-bit the damage was severe — AWQ w3 raised Qwen3-8B’s C4 perplexity from 10.4 to 23.8.
  • GPTQ degrades more gracefully at very low bits than the other methods, keeping some functionality even at 2-bit — but “some functionality” still means significant degradation versus full precision, not a usable daily driver.
  • Small models fall off the cliff earlier. Qwen3-14B lost about 1% MMLU under 4-bit GPTQ; Qwen3-0.6B lost around 10% under the same setting. The smaller the model, the less redundancy it has to absorb quantization noise — consistent with the long-tail-knowledge explanation above.

For vLLM users running AWQ or GPTQ in production, the operational takeaway is the same as for GGUF: 4-bit is the floor, and the format choice matters less than staying above it. Our GPTQ vs AWQ vs GGUF production guide covers throughput and serving differences between the formats.

Does this generalize beyond Qwen?

Partially verified, partially extrapolated — and the distinction matters. What’s verified: the same arXiv study compared Qwen3 against LLaMA3 under identical quantization and found Qwen3 degrades more at 3-bit and below. One plausible reading, offered by the authors: Qwen3 is trained closer to saturation for its parameter count, leaving less slack in the weights, so aggressive rounding bites harder. If that’s right, the newest, most efficiently-trained open models — the ones self-hosters most want to run — are precisely the ones most sensitive to low-bit quantization. Related work on factual recall under quantization (“Through a Compressed Lens,” arXiv 2505.13963) also finds knowledge-recall damage that benchmark averages understate.

What’s extrapolated: nobody has run IKP-style tiered knowledge probes across Llama, Gemma, and Mistral families at 55 quants each. The cliff’s location (between 4-bit and 3-bit for 27B-class dense models) is well-supported for Qwen; assume the shape generalizes, but verify the location before betting on it for a different family — the KL-divergence command above takes one evening per model.

Reasoning, for what it’s worth, follows a similar pattern: an empirical study of quantized reasoning models (arXiv 2504.04823, April 2025) found 4-bit broadly safe for chain-of-reasoning tasks with real degradation below that. Knowledge is not uniquely fragile — it’s just where the non-linearity was first measured cleanly.

When should you pay the VRAM cost for Q6 or Q8?

For most self-hosters: don’t — spend that VRAM on context or a bigger model, because the study found no measurable factual-recall gain above 5-bit. The cases where Q6_K or Q8_0 still earn their footprint:

  • No-RAG factual work in niche domains. If the model’s own weights are your knowledge base (medicine, law, niche history Q&A without retrieval), you are living in IKP’s upper tiers, where every bit of divergence from BF16 costs you. Q5_K_M minimum, Q6_K if it fits.
  • You can’t A/B test your workload. Q6_K at ~22 GB is cheap insurance on a 32 GB card when you have no way to measure what Q4 costs your tasks.
  • Coding backends are a judgment call. Code generation leans on both reasoning (quant-tolerant at 4-bit) and API recall (long-tail knowledge, quant-sensitive). If your local model backs Cline or a similar agent, hallucinated method names at Q4 that disappear at Q6 are this effect in the wild.

And the inverse: if you run RAG, quantization anxiety mostly evaporates. Retrieval puts the facts in the context window, so the model needs reasoning (intact at Q4) more than recall. A 27B Q4_K_M with a good retriever beats a 27B Q6_K without one for document work, and frees 5 GB of VRAM for embeddings.

When NOT to act on this study

  • Don’t generalize tier boundaries to MoE models. Qwen3.6 27B is dense. Mixture-of-experts models distribute knowledge differently across experts, and per-expert quantization sensitivity is not characterized by this data.
  • Don’t treat IKP as a proxy for your workload. It measures stored-fact recall. If you do summarization, translation, or RAG, your quality floor may sit lower than the knowledge cliff suggests.
  • Don’t assume two quants with the same label are equal. The 55-quant spread showed maker-to-maker variation at the same bit level; calibration data matters. When in doubt, measure KLD.
  • Don’t buy hardware off this article’s file sizes alone. They are approximate for 27B dense models; your model’s GGUF page lists exact sizes, and KV cache comes on top. For what the VRAM tiers cost in actual 2026 street prices, see the quantization quality-loss hardware guide on runaihome.com.

If you’re newer to quant formats and K-quant naming, our quantization guide for LLMs covers the vocabulary this article assumes.

FAQ

Is Q4_K_M still safe in 2026? Yes, for 14B+ dense models. In the Qwen3.6 27B study, 4-bit quants sat just above the knowledge cliff, and the follow-up on Qwen3.8 27B confirmed 4-bit holds up. Below 14B parameters, be more conservative — the arXiv Qwen3 study measured ~10% MMLU loss at 4-bit on a 0.6B model versus ~1% on 14B.

Does a higher quant fix hallucination? Only one kind. Quantization below 4-bit makes a model lose long-tail facts while staying fluent — it will confidently produce wrong answers it once knew. Moving from Q2/Q3 to Q4/Q5 recovers that. Hallucination that exists in the BF16 model won’t improve with more bits.

Should I run a 27B at Q3 or a 13B at Q5 on a 16 GB card? The 13B at Q5. The parameter advantage of the 27B is roughly linear, but the Q3 penalty is non-linear and concentrated in exactly the knowledge the extra parameters would have carried. Stay above the cliff first; maximize parameters second.

Sources

Prices are September 2026 street prices (not MSRP), verified against sold listings:

Your situationThe cardPriceWhere
27B at Q4_K_M–Q5_K_M, best $/GB of VRAMUsed RTX 3090 24GB~$1,150–$1,350Check price
27B at Q6_K with context headroomRTX 5090 32GB$3,822+ (MSRP $1,999 — never sold at it)Check price
Test Q4 vs Q6 on your workload before buyingRented 3090/4090from $0.07/hrVast.ai

Was this article helpful?

What self-hosting actually costs

Real cost breakdowns for self-hosted AI: hardware floors, power, maintenance hours, and the honest comparison against paying for it. No spam, unsubscribe anytime.