AI Quantization Explained: What Q4_K_M Means, Why 1-Bit Models Are Useless, and Which File You Should Actually Download
What's in this guide
If you followed my guide to running AI on your own computer, you have already hit this wall. You search for a model in LM Studio or on Hugging Face, and instead of one download button you get thirty files with names like UD-Q4_K_M, UD-Q2_K_XL and UD-IQ1_S. They range from 6GB to 55GB. Nothing explains which one you want.
Those names describe quantization, the most important thing to understand about local AI after how much memory you have. It is also having a moment: a benchmark of Qwen3.8 27B at every compression level hit the front page of Hacker News this week, and the local AI crowd on Reddit has spent the summer arguing about "brain damaged" 1-bit models. So here is the plain English version: what quantization is, how to read the file names, where quality falls off a cliff, and a simple rule for which file to download.
The short answer
Quantization is compression for AI models. A 4-bit version is roughly a third the size of the original and, on the benchmarks that matter, performs the same. A 2-bit version is noticeably worse but still usable for simple jobs. A 1-bit version is a toy: on a graduate-level science test it scores at the level of random guessing.
The rule: download the 4-bit file (the one with Q4_K_M in the name) whenever it fits in your memory with a few gigabytes to spare. If it does not fit, pick a smaller model at 4-bit rather than the same model at 2-bit or 1-bit. A 4-bit 9B model will beat a 1-bit 27B model every time, and it will be faster.
What quantization actually is
An AI model is a giant list of numbers. Qwen3.8 27B has 27 billion of them. In the original file each number is stored with 16 bits of precision, which works out to about 55GB on disk. That is more than most people have in RAM, let alone on a graphics card.
Quantization rounds those numbers off. Instead of storing each one to sixteen decimal places, so to speak, you store it to four. The file shrinks to roughly a quarter of the size, and the model runs faster because your computer is moving a quarter as much data around.
The obvious question is whether rounding off 27 billion numbers wrecks the model. Mostly it does not, because no single number matters much. What matters is the pattern across billions of them, and the pattern survives 4-bit rounding almost perfectly. Push the compression far enough, though, and the pattern breaks. That is the cliff we will get to in a minute.
How to read those file names
Here is the naming scheme used by GGUF files, which is the format LM Studio, Ollama and llama.cpp all use. Once you can read it, the wall of thirty files becomes a short menu.
Q4, Q8, IQ2, IQ1. The number is the bits per weight, and it is the only part that really matters. Q8 is 8-bit, half the size of the original with no measurable loss. Q4 is 4-bit, about a quarter of the size. Q2 and IQ1 are 2-bit and 1-bit. The "I" in IQ marks a method used at very low bit counts. Treat IQ2 as a slightly better 2-bit and IQ1 as 1-bit.
K_S, K_M, K_L, K_XL. Variants of the same bit count. S, M, L and XL mean small, medium, large and extra large: a few of the most sensitive parts of the model are kept at higher precision. Q4_K_M is the medium 4-bit and it is the default choice almost everywhere. Ollama downloads it when you do not specify anything else.
UD. Unsloth Dynamic, where the compression level varies layer by layer instead of being the same everywhere. Unsloth's Dynamic 3.0 files for Qwen3.8 27B came out in August and claim more than 10 percent better accuracy than other providers at the same file size. If you see a UD version and a plain version at the same bit count, take the UD one.
BF16 and F16. The uncompressed original. You do not want this unless you have a workstation with a very large graphics card and a specific reason.
To make this concrete, here are the actual sizes of the Unsloth Qwen3.8 27B files on Hugging Face as of September 9, 2026: BF16 is 54.7GB, Q8_0 is 29GB, UD-Q4_K_M is 16.5GB, UD-Q2_K_XL is 9.83GB and UD-IQ1_S is 6.19GB. Same model, nine times smaller at the bottom of that list.
The 4-bit sweet spot and the 1-bit cliff
Until recently the argument about how much quality you lose was mostly vibes. That changed with a benchmark from Piotr MigdaĆ at Quesma, published in late August and on the Hacker News front page on September 8. He ran the Unsloth versions of Qwen3.8 27B through three real tests: GPQA Diamond (graduate-level science questions), IFBench (does the model follow instructions precisely) and Terminal-Bench 2.1 (can it complete real programming jobs on its own). It cost him about $3,000 in rented GPU time, which is why nobody had done it before.
The findings, in order of how much they should change what you download:
4-bit matched the original. On Terminal-Bench, the hardest of the three, the 17GB Q4_K_M file scored the same as the 55GB original. On GPQA Diamond the difference was inside the noise. That puts the "quantized models are dumber" complaint to bed at 4-bit and above.
2-bit dropped a little and worked harder. The 2-bit UD-Q2_K_XL file was slightly worse on GPQA, noticeably worse on the programming test, and wrote about 25 percent more words to reach the same answers. On the instruction-following test, though, it was indistinguishable from the original. Simple jobs, fine. Hard jobs, you feel it.
1-bit collapsed. Both 1-bit files scored around random guessing on GPQA Diamond. Worse, turning up the model's reasoning effort made them score lower, because they would think until they ran out of room and return nothing. Unsloth's own page says the 1-bit file keeps about 72 percent of the original's word-by-word predictions. That number sounds fine. It is not, because the missing 28 percent is where the actual thinking lives.
The shape of that curve is the takeaway. Quality does not decline gently as the file shrinks. It stays flat, dips a little, then falls off a cliff between 2-bit and 1-bit. The 1-bit file exists so people can say a 27B model runs on a laptop. It does run. It just does not think.
What happened when I ran three versions myself
Benchmarks are one thing. I wanted to see whether I could feel the difference, so I downloaded the 4-bit, 2-bit and 1-bit Unsloth files of Qwen3.8 27B and ran each one in LM Studio on my 32GB desktop, the same machine from my RAM testing guide. No dedicated graphics card, so everything ran on the processor and system memory.
The 4-bit file loaded with room to spare and generated at a steady walking pace, a little under 5 words a second. Slow compared with ChatGPT, but the answers were the best I have had from anything on my own machine. I gave it a messy spreadsheet formula to fix and a three-paragraph email to rewrite. Both were right first time.
The 2-bit file was about a third faster and the email rewrite was just as good. The formula fix was wrong on the first attempt and right on the second, and it rambled more, exactly as the benchmark said it would.
The 1-bit file was the eye-opener. It typed fast and the output looked like English until you read it. The email rewrite lost the point of the email. The formula fix was confident nonsense. I asked it a simple factual question about a city I know well and it invented a river. I deleted the file after twenty minutes.
If you have 16GB of RAM rather than 32GB, none of the 27B files is a good fit. The 4-bit version alone is 16.5GB before the model has any room to work. That is not a reason to reach for the 1-bit file. It is a reason to pick a smaller model, which brings us to the chart.
Which file to download for your memory
Find your memory on the left, and download the file on the right. On a Mac, "memory" means unified memory. On a PC without a dedicated graphics card, it means system RAM. If you have a graphics card, use its VRAM figure instead, because the model runs far faster if it fits entirely on the card.
8GB. A 4-bit model in the 3 to 4 billion parameter range, about 2 to 3GB on disk. This is the same class of model that runs on a phone, and my offline phone AI guide covers what those are good for.
16GB. A 4-bit model in the 7 to 9 billion range, around 5 to 6GB on disk. This is the mainstream local AI experience and it is decent. You could squeeze a 4-bit 12 to 14B model in, but leave the browser closed.
24GB, or a 16GB graphics card. A 4-bit 12 to 14B model comfortably, or a 4-bit 27B model if you are willing to close everything else. A 24GB graphics card such as an RTX 4090 runs the 4-bit 27B file with room for a long conversation, which is exactly the setup the Quesma benchmark recommends.
32GB and up. The 4-bit 27B file, or an 8-bit 12 to 14B model if you prefer. In my experience the bigger model at 4-bit wins.
The rule behind the chart: file size plus 2 to 4GB of working room has to fit in your memory with the operating system still running. If it does not, step down a model size, never a bit count below 4.
LM Studio shows a small badge next to each file telling you whether it will fit on your machine, and it is usually right. Ollama simply picks Q4_K_M for you, which is the right call for almost everyone.
Two settings that matter more than the quant
Here is the part the file-name arguments miss. Once you are at 4-bit, two other settings change the answers you get far more than dropping to 8-bit would improve them.
Reasoning effort. Qwen3.8 27B ships with its thinking set to the highest level by default, and the benchmark found this mattered enormously: the best scores needed around 8,000 tokens (roughly 6,000 words) of hidden reasoning per question. On a machine without a graphics card, that is minutes of waiting. Turn it down to medium for everyday questions and save the high setting for problems that deserve it.
Context length. How much of the conversation the model can see. It eats memory on top of the file: for Qwen3.8 27B, about 2.3GB for every 32,000 tokens (roughly 24,000 words). If your model loads but crashes ten messages in, this is why. Cut the context length in half before you cut the quantization.
This is why two people can download the identical file and have completely different experiences. One has a graphics card and the defaults. The other is waiting three minutes per answer on a laptop with the effort set to maximum.
The verdict
Quantization sounds like an engineering detail and it is actually the deciding factor in whether local AI is worth your time. The good news from this week's benchmark is that the file most tools pick for you, the 4-bit Q4_K_M, is the right one. It performs like the original at a third of the size.
The rule to remember: 4-bit if it fits, a smaller model if it does not, and never 1-bit. Those tiny files are impressive engineering and they let a 27B model technically start on almost any laptop. But technically starting is not the same as thinking, and the 1-bit version does not think.
If you have not set up local AI yet, start with my guide to running AI on your own computer. And if you want to know where these files come from, my Hugging Face explainer covers the site every one of them is downloaded from.
Comments
Post a Comment