Last week I started a series about running open weights models on your own hardware. The goal was to give my subscribers a good understanding of what it is to run local, private AI on an entry level computer. The first post went really well, with more than 3000 views so far and dozens of hours watched. So today I’m pushing things forward, with yet another test, Google Gemma 4B E4B.
This one has just 4 billion parameters and file size is roughly 5GB, so it fits way better than the Ternary Bonsai we tested in the first video. For a quick reminder about what amount of RAM you need to comfortably run open weights models, you can have a look at the episode 4 of my series about how to choose an open weights model.
For this test I chose 2 tasks. First one is my usual “what is a dolphin?” question, but second one is a simple code generation task. Google Gemma 4 E4B managed both very well. In terms of time, it was usable, the question was answered in a few seconds, while the code generation took about 2.5 minutes. Kind reminder that all recordings in these videos are not edited in any way, no sped up, nothing, so you can have a real understanding of what it means to run the tested model.
All in all, Google Gemma 4 E4B is a nice surprise. I would not use it for very intense code generation tasks, but for quick functions, or summarization, I think it’s an interesting choice. The fact that it only takes 5GB of memory, leaving a lot of room for other programs on my tiny 16GB M1 MacBookPro is a big plus too.
If you have any suggestions about which models should I test next, drop a comment on that video. An of course, if you like this kind of content: like, share and subscribe!
Thanks for watching and spreading the word, you’re helping local, sovereign AI to come a little bit closer with every share.
Top comments (0)