DEV Community

#llm

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
What running an LLM in production actually costs you

What running an LLM in production actually costs you

1
Comments 1
5 min read
How OpenAI's Models Escaped a Sandbox and Breached Hugging Face

How OpenAI's Models Escaped a Sandbox and Breached Hugging Face

2
Comments
10 min read
More Context Made My Classifier Worse: Building a Machine-Maintained Failure Taxonomy

More Context Made My Classifier Worse: Building a Machine-Maintained Failure Taxonomy

Comments
4 min read
How I cut AI generation errors in half without burning more LLM calls

How I cut AI generation errors in half without burning more LLM calls

Comments
3 min read
AI Agent Sandboxing: Contain the Blast Radius

AI Agent Sandboxing: Contain the Blast Radius

2
Comments 2
9 min read
Speculative Decoding: Faster On-Device LLMs

Speculative Decoding: Faster On-Device LLMs

1
Comments 4
6 min read
Your AI coding agent gets expensive one reasonable decision at a time

Your AI coding agent gets expensive one reasonable decision at a time

Comments
3 min read
RAG Without Vectors: How LLMs Are Learning to Navigate Documents Like Humans

RAG Without Vectors: How LLMs Are Learning to Navigate Documents Like Humans

Comments
6 min read
How I Built a 6-Agent AI Recruitment System on Qwen Cloud in less than 28 Days

How I Built a 6-Agent AI Recruitment System on Qwen Cloud in less than 28 Days

Comments
1 min read
Chunked Prefill: Why One Long Prompt Freezes Your LLM Server

Chunked Prefill: Why One Long Prompt Freezes Your LLM Server

Comments 1
7 min read
I Ran 1,000 LLM Evals Over 12 Months. Here's What Actually Moved the Needle

I Ran 1,000 LLM Evals Over 12 Months. Here's What Actually Moved the Needle

Comments
4 min read
[AI] Optimizing vLLM Serving: AWQ, GPTQ, & GGUF | SLM Playbook

[AI] Optimizing vLLM Serving: AWQ, GPTQ, & GGUF | SLM Playbook

Comments
4 min read
We let an AI agent hit a database 1034 times. Text-to-SQL ran 23 unsafe ops. The policy layer ran zero

We let an AI agent hit a database 1034 times. Text-to-SQL ran 23 unsafe ops. The policy layer ran zero

Comments
4 min read
What Really Concerns Me is One of the Biggest Issues with AI Coding Agents: Context Isolation and Task Coordination

What Really Concerns Me is One of the Biggest Issues with AI Coding Agents: Context Isolation and Task Coordination

Comments
6 min read
RAG en entreprise : ce que j'ai appris en le déployant chez un industriel (REX)

RAG en entreprise : ce que j'ai appris en le déployant chez un industriel (REX)

Comments
4 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.