Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
llamacpp
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
Gemma 4 at Over 70 Tokens/s on a 2021 Laptop's 4 GB GPU: The Live Demo, Step by Step
xbill
xbill
xbill
Follow
for
Google Developer Experts
Oct 2
Gemma 4 at Over 70 Tokens/s on a 2021 Laptop's 4 GB GPU: The Live Demo, Step by Step
#
gemma
#
llamacpp
#
cuda
#
huggingface
9
 reactions
Comments
Add Comment
12 min read
Inside My llama.cpp Setup: Tuning Qwen 3.8 27B for 512K Context
Dmitry Amelchenko
Dmitry Amelchenko
Dmitry Amelchenko
Follow
Oct 3
Inside My llama.cpp Setup: Tuning Qwen 3.8 27B for 512K Context
#
ai
#
llm
#
localllm
#
llamacpp
Comments
Add Comment
7 min read
llama-server's sleep mode loses or crashes on a request that arrives just before it sleeps
The Homelab Postmortem
The Homelab Postmortem
The Homelab Postmortem
Follow
Oct 3
llama-server's sleep mode loses or crashes on a request that arrives just before it sleeps
#
llm
#
llamacpp
#
selfhosted
#
debugging
1
 reaction
Comments
1
 comment
7 min read
When Chat Templates Go Wrong
pielouNW
pielouNW
pielouNW
Follow
Oct 1
When Chat Templates Go Wrong
#
llm
#
opensource
#
ai
#
llamacpp
Comments
Add Comment
7 min read
llama.cpp vs Ollama in 2026: Which Runtime Should You Run?
Rost
Rost
Rost
Follow
Sep 14
llama.cpp vs Ollama in 2026: Which Runtime Should You Run?
#
llamacpp
#
ollama
#
llm
#
selfhosting
Comments
Add Comment
18 min read
A 4 GB Laptop GPU vs a 6-Core CPU on Gemma 4, Re-Measured in ABBA Order: 4.1x
xbill
xbill
xbill
Follow
for
Google Developer Experts
Sep 23
A 4 GB Laptop GPU vs a 6-Core CPU on Gemma 4, Re-Measured in ABBA Order: 4.1x
#
gemma
#
llamacpp
#
cuda
#
benchmarking
10
 reactions
Comments
2
 comments
11 min read
ROCm vs Vulkan for AMD Local LLM Hosting: 2026 Guide
Rost
Rost
Rost
Follow
Sep 12
ROCm vs Vulkan for AMD Local LLM Hosting: 2026 Guide
#
llm
#
selfhosting
#
llamacpp
#
ollama
Comments
2
 comments
20 min read
KV Cache on 16 GB GPUs: Making Long Context Actually Fit
Rost
Rost
Rost
Follow
Sep 11
KV Cache on 16 GB GPUs: Making Long Context Actually Fit
#
llm
#
llamacpp
#
vllm
#
ollama
Comments
1
 comment
22 min read
How to pick --n-cpu-moe in llama.cpp: Qwen3.6 35B-A3B on 12, 16 and 24 GB GPUs
Donald Lee
Donald Lee
Donald Lee
Follow
Sep 29
How to pick --n-cpu-moe in llama.cpp: Qwen3.6 35B-A3B on 12, 16 and 24 GB GPUs
#
llm
#
llamacpp
#
localai
#
gpu
1
 reaction
Comments
2
 comments
3 min read
Running llama.cpp on a 32 GB MacBook Air: A Direct Comparison with Ollama
Vishnu Hari Dadhich
Vishnu Hari Dadhich
Vishnu Hari Dadhich
Follow
Sep 7
Running llama.cpp on a 32 GB MacBook Air: A Direct Comparison with Ollama
#
ai
#
localllm
#
llamacpp
#
ollama
Comments
1
 comment
9 min read
Fix Local LLM Quality: Context Stacking & Rope Freq Tweaks
Umair Bilal
Umair Bilal
Umair Bilal
Follow
Aug 23
Fix Local LLM Quality: Context Stacking & Rope Freq Tweaks
#
localllms
#
aiagents
#
ollama
#
llamacpp
Comments
Add Comment
8 min read
Ollama vs LM Studio 2026: Ollama Wins for Devs
Shaam
Shaam
Shaam
Follow
Sep 25
Ollama vs LM Studio 2026: Ollama Wins for Devs
#
ollama
#
lmstudio
#
localllm
#
llamacpp
Comments
1
 comment
7 min read
"V cache quantization requires flash_attn" — the llama.cpp error that quietly halves your context window
Jasur Yuldoshev
Jasur Yuldoshev
Jasur Yuldoshev
Follow
Aug 21
"V cache quantization requires flash_attn" — the llama.cpp error that quietly halves your context window
#
llamacpp
#
llm
#
performance
#
debugging
Comments
Add Comment
10 min read
llama.cpp vs Ollama: Which Should You Run in 2026?
Mr Say Nothing
Mr Say Nothing
Mr Say Nothing
Follow
Sep 18
llama.cpp vs Ollama: Which Should You Run in 2026?
#
ai
#
llamacpp
#
ollama
#
localllm
Comments
1
 comment
6 min read
How to Run GGUF Models Locally: Ollama, llama.cpp & vLLM
Mr Say Nothing
Mr Say Nothing
Mr Say Nothing
Follow
Sep 18
How to Run GGUF Models Locally: Ollama, llama.cpp & vLLM
#
ai
#
gguf
#
ollama
#
llamacpp
Comments
1
 comment
5 min read
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account