DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Why Your AI App Is Failing in Production (And It’s Not an ML Problem)

Why Your AI App Is Failing in Production (And It’s Not an ML Problem)

Comments
3 min read
I Thought I Used 1 Billion Tokens in Claude Code Last Week. Turns Out 97% Was Cache.

I Thought I Used 1 Billion Tokens in Claude Code Last Week. Turns Out 97% Was Cache.

Comments
4 min read
Orchestrating 5 Agents in LangGraph Without Deadlocks

Orchestrating 5 Agents in LangGraph Without Deadlocks

Comments 1
5 min read
Qwen3 4B vs 8B vs 14B for Writing Correction: 60 Local Ollama Responses on Windows

Qwen3 4B vs 8B vs 14B for Writing Correction: 60 Local Ollama Responses on Windows

Comments
5 min read
Qwen Code vs Aider vs OpenCode: I Ran the Same 12 Tasks Through 3 Local CLI Agents on One RTX 4070

Qwen Code vs Aider vs OpenCode: I Ran the Same 12 Tasks Through 3 Local CLI Agents on One RTX 4070

Comments
6 min read
Your Intel Laptop Can Run 30B Models Now. No NVIDIA. No Cloud. No Problem.

Your Intel Laptop Can Run 30B Models Now. No NVIDIA. No Cloud. No Problem.

Comments
11 min read
Data Gravity: The Real Cost of API-First AI

Data Gravity: The Real Cost of API-First AI

Comments
9 min read
vLLM reinvented the operating system, and nobody told you

vLLM reinvented the operating system, and nobody told you

1
Comments
12 min read
Using Multiple LLM Providers with the Laravel AI SDK

Using Multiple LLM Providers with the Laravel AI SDK

Comments
8 min read
Running local and cloud models in the same coding agent: what actually ships in 2026

Running local and cloud models in the same coding agent: what actually ships in 2026

Comments
4 min read
How I Built a $0 LLM Production Stack with 46 Free APIs

How I Built a $0 LLM Production Stack with 46 Free APIs

Comments
3 min read
Node.js vLLM LLM Inference: 50ms Latency on RTX 4090

Node.js vLLM LLM Inference: 50ms Latency on RTX 4090

Comments
8 min read
I measured his app with his own code. He measured my claim with his own corpus.

I measured his app with his own code. He measured my claim with his own corpus.

1
Comments
10 min read
I pre-registered a prediction that my own finding would fail on this market. The product list held; the advice did not.

I pre-registered a prediction that my own finding would fail on this market. The product list held; the advice did not.

Comments
7 min read
RAG Database with VLM as Extractor

RAG Database with VLM as Extractor

Comments
3 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.