DEV Community

Nerav Doshi
Nerav Doshi

Posted on Edited on Originally published at pipelineandprompts.com

Compared Quantization Levels: Q4 vs Q8 vs FP16 on llama3.2:1b

A model's weights are normally stored at high precision — FP16, 16-bit floating point. Quantization compresses those numbers down to fewer bits to shrink the file and memory footprint, trading some precision for space. Same idea as squashing a high-res photo down to a smaller JPEG: the file shrinks a lot, technically some detail is lost, and in practice you usually can't tell.

Before pulling anything new for this one, I ran ollama show llama3.2:1b out of curiosity — and immediately had to correct myself. The model I'd been calling "the default" in Entries 01 and 03 turned out to already be Q8_0, not some unspecified baseline. Worth owning that plainly rather than quietly fixing it: I'd been comparing against a number without actually knowing what it was.

So, three levels this time: the Q8 I already had, an explicit q4_K_M build, and a full-precision fp16 pull. Loaded each one, sent the same question, grabbed ollama ps and ps aux for all three.

Q4_K_M Q8_0 (the "default" from Entry 01) FP16
Download size 807 MB 1.3 GB 2.5 GB
ollama ps loaded size 997 MB 1.5 GB 2.7 GB
Process RSS ~0.94 GB ~1.24 GB ~2.57 GB

The memory numbers lined up almost too neatly — roughly doubling at each step, right along the bit-width math. Cleaner relationship than I expected going in; quantization level turns out to be a genuinely predictable lever for sizing memory.

Quality was a different story, and honestly the more interesting one. Same question to all three: "what are good free AI tools that can create simple Google Slides using instructions?" None of them gave a clean answer. All three padded a decent core suggestion (Canva) with tools that have nothing to do with slides — Q4 threw in Midjourney and DALL-E, FP16 added Artbreeder, Deep Dream Generator, Prisma. Going all the way up to full precision didn't clean any of that up. Q4 wasn't noticeably worse than FP16 here, which is not what I expected walking in.

$ ollama show llama3.2:1b
  quantization        Q8_0

$ ollama show llama3.2:1b-instruct-q4_K_M
  quantization        Q4_K_M

$ ollama ps
NAME                         ID              SIZE      PROCESSOR    CONTEXT    UNTIL
llama3.2:1b-instruct-fp16    2887c3d03e74    2.7 GB    100% GPU     4096       4 minutes from now

$ ps aux | grep ollama
flyers  10763  0.3  15.7  438101456  2630128  ??  S  llama-server --model ... -c 4096
flyers   2256  0.0   0.3  436826160    56448  ??  S  ollama serve

$ ollama ps
NAME                           ID              SIZE      PROCESSOR    CONTEXT    UNTIL
llama3.2:1b-instruct-q4_K_M    22bc6b92eb01    997 MB    100% GPU     4096       4 minutes from now

$ ps aux | grep ollama
flyers  10786  0.2  5.7  436484688  963808  ??  S  llama-server --model ... -c 4096
flyers   2256  0.0  0.4  436826160   60400  ??  S  ollama serve
Enter fullscreen mode Exit fullscreen mode

I don't think this proves quantization is "safe" in some general sense — it was one soft, open-ended question, not a real eval. My guess is the actual quality gap shows up on harder stuff: math, code, anything that needs precise instruction-following rather than a casual recommendation list. But the memory story held up cleanly, and now I've got a habit I didn't have before: check ollama show <model> before assuming you know what precision a "default" pull actually gave you. I clearly didn't, back in Entry 01.

Top comments (0)