DEV Community

#llm

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
Stop explaining yourself to Claude

Stop explaining yourself to Claude

Comments
3 min read
Multi-Turn Email Conversations for LLM Agents

Multi-Turn Email Conversations for LLM Agents

1
Comments
4 min read
Shipping an LLM-powered "lie detector" in a Flutter dating app

Shipping an LLM-powered "lie detector" in a Flutter dating app

Comments
3 min read
The best bug reports were written by the suspect

The best bug reports were written by the suspect

1
Comments 1
1 min read
From Chatbot to Mailbox: Persistent Agent Memory in Threads

From Chatbot to Mailbox: Persistent Agent Memory in Threads

1
Comments
5 min read
Sampling strategies compared: temperature, top-p, top-k, min-p, and what actually works in production

Sampling strategies compared: temperature, top-p, top-k, min-p, and what actually works in production

Comments
9 min read
Why your AI Agent needs a sandbox, not a blank check 🛡️

Why your AI Agent needs a sandbox, not a blank check 🛡️

Comments
1 min read
I tested 3 models as AI agent quality inspectors: the stronger the model, the more valid work it rejects

Precision-recall and the false positive mirage

I tested 3 models as AI agent quality inspectors: the stronger the model, the more valid work it rejects

10
Comments 60
5 min read
I Added More Verifiers to Catch Bad AI Findings. Precision Stopped Improving.

I Added More Verifiers to Catch Bad AI Findings. Precision Stopped Improving.

Comments
3 min read
SQL AI Database Solutions: Building a Safe Text-to-SQL App with Streamlit and Hugging Face

SQL AI Database Solutions: Building a Safe Text-to-SQL App with Streamlit and Hugging Face

38
Comments 2
5 min read
We create a way to unload Qwen2.5 KV cache to RAM.

We create a way to unload Qwen2.5 KV cache to RAM.

Comments
3 min read
I catalogued 32 real AI-agent failures, then marked the ones we cannot stop

I catalogued 32 real AI-agent failures, then marked the ones we cannot stop

1
Comments 2
4 min read
LLM Evals For Developer Tools: Useful, Correct, Safe

Risk of agents gaming tests to pass

LLM Evals For Developer Tools: Useful, Correct, Safe

52
Comments 30
18 min read
# Agentic Systems for Big Query Handling in Distributed Environments: The Complete Engineering Guide

# Agentic Systems for Big Query Handling in Distributed Environments: The Complete Engineering Guide

Comments
20 min read
Introducing Duplex: A Zero-Backend, Multiplexed LLM Inference Engine for True Client-Side Parallel AI

Introducing Duplex: A Zero-Backend, Multiplexed LLM Inference Engine for True Client-Side Parallel AI

Comments
4 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.