
Llama 3.1 8B in 16-bit precision weighs 14.96 GiB. Squash it to Q4_K_M and it weighs 4.58 GiB, generates tokens roughly 2.5 times faster on the same machine, and, on the best available evidence, loses you about a rounding error of quality. That is the whole pitch of quantisation, and it is why the model you thought needed a server rack now runs on the laptop you already own.
This is the next instalment of our AI on a Budget series, and it is the one that makes the local-model advice actually work. Every guide says “grab a 4-bit GGUF” as though the letters were self-explanatory. They are not, and picking the wrong format costs you speed, quality or both.
What quantisation actually does
A model is a giant pile of numbers (weights). Stored at 16 bits each, an 8-billion-parameter model needs about 16 GB before it has processed a single token. Quantisation stores each weight in fewer bits, usually 4 to 8, plus a bit of bookkeeping so the maths stays close to the original. Less memory to hold, less memory to read per token, and since local inference is mostly limited by memory bandwidth, faster generation as well.
The price is rounding error. Whether you notice it depends on the bit count, the method and the task. Maths and code suffer first; commonsense and chat barely move.
The real numbers for an 8B model
The llama.cpp project publishes its own table for Llama 3.1 8B. Sizes and generation speeds (tokens per second at 128 tokens) from that table:
- F16: 14.96 GiB, 29.17 t/s
- Q8_0: 7.95 GiB, 50.93 t/s
- Q6_K: 6.14 GiB, 58.67 t/s
- Q5_K_M: 5.33 GiB, 67.23 t/s
- Q4_K_M: 4.58 GiB, 71.93 t/s
- Q3_K_M: 3.74 GiB, 71.68 t/s
- Q2_K: 2.95 GiB, 79.85 t/s
Note the shape. Going from 16-bit to Q4_K_M cuts size by about 70 percent. Going from Q4_K_M to Q3_K_M saves another 0.84 GiB and gains no speed at all. Below 4 bits you are paying quality for very little.
How much quality do you lose?
A January 2026 evaluation of llama.cpp quants on Llama-3.1-8B-Instruct measured benchmark averages against F16. Q5_0 landed at +0.65 percent (noise), Q4_K_S at -0.43 percent, Q3_K_L at -0.99 percent, and Q3_K_S at -5.73 percent. WikiText-2 perplexity went from 7.32 at F16 to a range of 7.33 to 8.96 across formats. Maths reasoning (GSM8K) was the most fragile task; HellaSwag barely flinched.
The authors call Q5_0 the accuracy pick and Q4_K_S the balanced default. Q4_K_M sits a touch above Q4_K_S in size and is the safe community default. If your workload is arithmetic-heavy, step up to Q5_K_M or Q6_K rather than trusting a 3-bit file.
GGUF, AWQ, GPTQ: which one and where
They are three answers to three different situations.
- GGUF is the llama.cpp format. It runs on CPU, Apple Silicon and GPUs, and can split a model between RAM and VRAM. If you are on a laptop or a Mac, this is your only real option, and it works with Ollama and LM Studio out of the box.
- AWQ (activation-aware weight quantisation, MLSys 2024 best paper) protects the roughly 1 percent of weights that matter most, identified by looking at activations. The authors report more than 3x speedup over the Hugging Face FP16 implementation on GPUs. It is a GPU format for serving.
- GPTQ is the older GPU format. Still widely available, but the newer tests are less kind to it.
One benchmark on an H200 with Qwen2.5-32B under vLLM is worth reading twice. Perplexity: FP16 6.56, GGUF Q4_K_M 6.74, AWQ 6.84, GPTQ 6.90. HumanEval pass@1: FP16 56.1 percent, GGUF and AWQ both 51.8 percent, GPTQ 46.3 percent. Code generation is where GPTQ gave up about ten points.
The trap: the kernel matters more than the format
The same benchmark measured AWQ at 68 tokens per second with the default kernel and 741 with the Marlin kernel. GPTQ went from 277 to 712. Swapping the kernel under identical weights moved throughput by up to 10.9x, which dwarfs any difference between the algorithms. At the quoted H200 rate of $4.80 an hour, that spread turned into roughly $1.80 per million tokens for AWQ with Marlin against $19.61 for AWQ with the default kernel. If you self-host with vLLM, check that your build uses the fast kernel before you blame the format.
Also, do not budget VRAM from file size. Weights are only part of it: the KV cache grows with context length, and in one RTX 5090 test total memory use converged to within about 1 GB regardless of format. Leave 20 to 30 percent headroom.
What this means
- Default to Q4_K_M for chat, summarising and drafting. Move to Q5_K_M or Q6_K for maths and code.
- Skip anything under 4 bits unless memory is the only constraint. The speed gain is close to nil.
- Laptop or Mac: GGUF. Single-GPU server on vLLM: AWQ, with the Marlin kernel confirmed.
- Budget: an 8B model at Q4_K_M needs about 4.6 GiB for weights, so an 8 GB card or a 16 GB laptop is enough. That is hardware you already own, which turns a per-token API bill into a fixed cost of zero for the routine jobs.
- Keep an API model for the hard 10 percent. Route the rest locally, as covered in our routing guide.
Downsides worth stating: quantised models can fail differently from full-precision ones on long reasoning chains, benchmarks above come from third parties on specific hardware, and a 4-bit 8B model is still an 8B model. It will not out-think a frontier API.
Did you know: AWQ is scale-based and does not train anything, so quantising a model with it takes minutes of calibration on a small sample rather than a fine-tuning run.
Sources
- llama.cpp quantize README (size and speed table)
- Which Quantization Should I Use? Unified evaluation of llama.cpp quantization on Llama-3.1-8B-Instruct
- AWQ: Activation-aware Weight Quantization (arXiv)
- Spheron: GPTQ vs AWQ vs GGUF comparison
Related on Top Tool Stack: AI on a Budget: cut your AI costs · Prompt caching