DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Reasoning happens before the response

Reasoning happens before the response

Comments
5 min read
What building an LLM inference engine from scratch taught me about compiler design

What building an LLM inference engine from scratch taught me about compiler design

1
Comments 1
4 min read
Qwen 3.6 27B and 35B MTP vs Standard on 16GB GPU

Qwen 3.6 27B and 35B MTP vs Standard on 16GB GPU

Comments
8 min read
The ChatGPT Invisibility Bug: Why High-Quality Content Fails to Index in LLM Search

The ChatGPT Invisibility Bug: Why High-Quality Content Fails to Index in LLM Search

1
Comments
5 min read
AI Weekly — 2026-05-22 to 2026-05-29 | Anthropic's $965B Moment and the Infrastructure Bet

AI Weekly — 2026-05-22 to 2026-05-29 | Anthropic's $965B Moment and the Infrastructure Bet

Comments
5 min read
Stop letting LLMs hallucinate dates — a tool for AI agents

Stop letting LLMs hallucinate dates — a tool for AI agents

5
Comments 1
2 min read
How I built an intent drift detector for LLM agents

How I built an intent drift detector for LLM agents

Comments 3
1 min read
Continuous batching wrecked our p99 latency. Here's the trace.

Continuous batching wrecked our p99 latency. Here's the trace.

Comments
4 min read
Why Enterprise AI Needs Structured Dissent, Not Just More Agents

Why Enterprise AI Needs Structured Dissent, Not Just More Agents

Comments 1
6 min read
GLM-5.2 vs Claude Opus: What the Numbers Actually Say for Developers

GLM-5.2 vs Claude Opus: What the Numbers Actually Say for Developers

1
Comments
7 min read
Building a cost-efficient LLM caching layer in Python

Building a cost-efficient LLM caching layer in Python

Comments
5 min read
Why Claude Code Sessions Diverge: A Mechanism Catalog

Why Claude Code Sessions Diverge: A Mechanism Catalog

Comments
3 min read
Gemma4 Apex GGUF, Ollama Context Optimization, & Llama3 Benchmarks

Gemma4 Apex GGUF, Ollama Context Optimization, & Llama3 Benchmarks

Comments
3 min read
How Do You Fit a Trillion-Parameter Model Into a Kubernetes Cluster?

How Do You Fit a Trillion-Parameter Model Into a Kubernetes Cluster?

Comments
17 min read
Prompt Engineering: The Skill That Makes AI Work Better

Prompt Engineering: The Skill That Makes AI Work Better

Comments
2 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.