DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
vLLM vs llama.cpp vs Ollama: What Happens When Your Model Doesn't Fit in 24GB VRAM

vLLM vs llama.cpp vs Ollama: What Happens When Your Model Doesn't Fit in 24GB VRAM

Comments
6 min read
26 AI Models Compared: A 2026 Cost Guide (GPT-4o vs Claude vs DeepSeek vs Local)

26 AI Models Compared: A 2026 Cost Guide (GPT-4o vs Claude vs DeepSeek vs Local)

1
Comments
8 min read
NodeLLM 1.17: MCP Sampling, Concurrent Tool Execution, and Smarter ORM Control

NodeLLM 1.17: MCP Sampling, Concurrent Tool Execution, and Smarter ORM Control

Comments
4 min read
Summary — Your Next Steps as an AI Architect

Summary — Your Next Steps as an AI Architect

Comments
3 min read
Building an Autonomous Budget Gate: Optimizing LLM Costs with Speculative Runtime Execution

Building an Autonomous Budget Gate: Optimizing LLM Costs with Speculative Runtime Execution

Comments
3 min read
Ten Layers of AI Skill Construction: A Systematic Framework from Prompts to Business Closed Loops

Ten Layers of AI Skill Construction: A Systematic Framework from Prompts to Business Closed Loops

Comments
9 min read
Fable May Not Be the Best Choice for Some Engineers

Fable May Not Be the Best Choice for Some Engineers

Comments
4 min read
Your Guardrails Are a Firewall. Your Failures Are a Cascade

Your Guardrails Are a Firewall. Your Failures Are a Cascade

Comments
5 min read
"183 Local Tools, Zero Guardrails: What Local MCP Gets Wrong About 'Privacy'"

"183 Local Tools, Zero Guardrails: What Local MCP Gets Wrong About 'Privacy'"

Comments
3 min read
AI Governance — EU AI Act Compliance, Risk Assessment, and Audit Logging

AI Governance — EU AI Act Compliance, Risk Assessment, and Audit Logging

Comments
9 min read
LLM-as-Judge Is Too Lenient. Here's a Cheap Fix: Judge Refute (Maybe) Arbitrate

LLM-as-Judge Is Too Lenient. Here's a Cheap Fix: Judge Refute (Maybe) Arbitrate

1
Comments 1
8 min read
Tiered Context Loading: Fit a Huge Agent Registry in Your Context Window

Tiered Context Loading: Fit a Huge Agent Registry in Your Context Window

Comments
7 min read
The KV cache, why LLM inference is memory-bound, not compute-bound

The KV cache, why LLM inference is memory-bound, not compute-bound

Comments
4 min read
Vision Language Models — When AI Learns to See and Talk (Part 3 of 3)

Vision Language Models — When AI Learns to See and Talk (Part 3 of 3)

Comments
13 min read
How LLM Function Calling Actually Works — From Tokens to Tool Orchestration

How LLM Function Calling Actually Works — From Tokens to Tool Orchestration

Comments
7 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.