A model's weights are normally stored at high precision — FP16, 16-bit floating point. Quantization compresses those numbers down to fewer bits to shrink the file and memory footprint, trading some precision for space. Same idea as squashing a high-res photo down to a smaller JPEG: the file shrinks a lot, technically some detail is lost, and in practice you usually can't tell.
Before pulling anything new for this one, I ran ollama show llama3.2:1b out of curiosity — and immediately had to correct myself. The model I'd been calling "the default" in Entries 01 and 03 turned out to already be Q8_0, not some unspecified baseline. Worth owning that plainly rather than quietly fixing it: I'd been comparing against a number without actually knowing what it was.
So, three levels this time: the Q8 I already had, an explicit q4_K_M build, and a full-precision fp16 pull. Loaded each one, sent the same question, grabbed ollama ps and ps aux for all three.
| Q4_K_M | Q8_0 (the "default" from Entry 01) | FP16 | |
|---|---|---|---|
| Download size | 807 MB | 1.3 GB | 2.5 GB |
ollama ps loaded size |
997 MB | 1.5 GB | 2.7 GB |
| Process RSS | ~0.94 GB | ~1.24 GB | ~2.57 GB |
The memory numbers lined up almost too neatly — roughly doubling at each step, right along the bit-width math. Cleaner relationship than I expected going in; quantization level turns out to be a genuinely predictable lever for sizing memory.
Quality was a different story, and honestly the more interesting one. Same question to all three: "what are good free AI tools that can create simple Google Slides using instructions?" None of them gave a clean answer. All three padded a decent core suggestion (Canva) with tools that have nothing to do with slides — Q4 threw in Midjourney and DALL-E, FP16 added Artbreeder, Deep Dream Generator, Prisma. Going all the way up to full precision didn't clean any of that up. Q4 wasn't noticeably worse than FP16 here, which is not what I expected walking in.
$ ollama show llama3.2:1b
quantization Q8_0
$ ollama show llama3.2:1b-instruct-q4_K_M
quantization Q4_K_M
$ ollama ps
NAME ID SIZE PROCESSOR CONTEXT UNTIL
llama3.2:1b-instruct-fp16 2887c3d03e74 2.7 GB 100% GPU 4096 4 minutes from now
$ ps aux | grep ollama
flyers 10763 0.3 15.7 438101456 2630128 ?? S llama-server --model ... -c 4096
flyers 2256 0.0 0.3 436826160 56448 ?? S ollama serve
$ ollama ps
NAME ID SIZE PROCESSOR CONTEXT UNTIL
llama3.2:1b-instruct-q4_K_M 22bc6b92eb01 997 MB 100% GPU 4096 4 minutes from now
$ ps aux | grep ollama
flyers 10786 0.2 5.7 436484688 963808 ?? S llama-server --model ... -c 4096
flyers 2256 0.0 0.4 436826160 60400 ?? S ollama serve
I don't think this proves quantization is "safe" in some general sense — it was one soft, open-ended question, not a real eval. My guess is the actual quality gap shows up on harder stuff: math, code, anything that needs precise instruction-following rather than a casual recommendation list. But the memory story held up cleanly, and now I've got a habit I didn't have before: check ollama show <model> before assuming you know what precision a "default" pull actually gave you. I clearly didn't, back in Entry 01.
Top comments (0)