DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Achieving Maximum Throughput on vLLM with a Single RTX 3090: A Production Guide for 7B LLMs

Achieving Maximum Throughput on vLLM with a Single RTX 3090: A Production Guide for 7B LLMs

2
Comments 1
4 min read
Seu agente de IA está desperdiçando 13.000 tokens antes de dizer "oi"

Seu agente de IA está desperdiçando 13.000 tokens antes de dizer "oi"

Comments
4 min read
Why AI Hallucinates Even When It Knows the Answer

Why AI Hallucinates Even When It Knows the Answer

1
Comments
5 min read
The Production Agent Checklist: What Every AI Agent Needs Before It Touches Real Users

The Production Agent Checklist: What Every AI Agent Needs Before It Touches Real Users

Comments
9 min read
How I Built a Hallucination Detector for RAG Pipelines in Python

How I Built a Hallucination Detector for RAG Pipelines in Python

Comments 1
3 min read
Stop Hardcoding Model Fallbacks: Let Production Data Pick Your Paths

Stop Hardcoding Model Fallbacks: Let Production Data Pick Your Paths

Comments
8 min read
SEO Is Dead? No. But the Game Changed.

SEO Is Dead? No. But the Game Changed.

Comments
11 min read
The AI Engineer's Toolkit: Moving Beyond Prompt Engineering to Build Robust AI Applications

The AI Engineer's Toolkit: Moving Beyond Prompt Engineering to Build Robust AI Applications

1
Comments
5 min read
AI Agents in Production Are Flying Blind — AgentLens Fixes That

AI Agents in Production Are Flying Blind — AgentLens Fixes That

Comments 1
2 min read
Your AI agent wastes 13,000 tokens before saying "hello"

Your AI agent wastes 13,000 tokens before saying "hello"

1
Comments
4 min read
Your AI Agent Can Be Socially Engineered. Here Are 3 Attacks That Prove It.

Your AI Agent Can Be Socially Engineered. Here Are 3 Attacks That Prove It.

4
Comments
4 min read
Building a Context-Aware AI Chat Without a Vector Database

Building a Context-Aware AI Chat Without a Vector Database

Comments
6 min read
SimCore: I built a social simulation engine where LLM agents live on a real map of your city

SimCore: I built a social simulation engine where LLM agents live on a real map of your city

Comments
1 min read
Multi-Model LLM Orchestration with OpenRouter

Multi-Model LLM Orchestration with OpenRouter

Comments
6 min read
I Tried Speculative Decoding on RTX 4060 8GB — Every Config Was Slower Than Baseline

I Tried Speculative Decoding on RTX 4060 8GB — Every Config Was Slower Than Baseline

1
Comments
8 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.