DEV Community

#llmops

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
Stop Guessing Which Model Is Better: Amazon Bedrock Model Evaluation Hands-On

Stop Guessing Which Model Is Better: Amazon Bedrock Model Evaluation Hands-On

Comments
5 min read
The throttle that wasn't a cap: rate vs sum in agent budgets

The throttle that wasn't a cap: rate vs sum in agent budgets

Comments
5 min read
Invoked, not executed

Invoked, not executed

Comments
5 min read
Your AI Agent Doesn't Need a Bigger Context Window. It Needs an Eviction Policy.

Your AI Agent Doesn't Need a Bigger Context Window. It Needs an Eviction Policy.

2
Comments 2
5 min read
Switchboard: building a tool router so your AI agent stops drowning in MCP tools

Switchboard: building a tool router so your AI agent stops drowning in MCP tools

Comments
13 min read
Detailed instructions written for an earlier generation of AI models become harmful on today's models

Detailed instructions written for an earlier generation of AI models become harmful on today's models

Comments
7 min read
We caught our AI agent building backdoors to run itself more

We caught our AI agent building backdoors to run itself more

1
Comments 2
6 min read
The Dedicated-GPU Trap: Why "Frontier, Per-Token" Wins Until You Have a Dozen Clients

The Dedicated-GPU Trap: Why "Frontier, Per-Token" Wins Until You Have a Dozen Clients

Comments
5 min read
Fine-Tuning Is Often the Wrong First Move: Introducing Harneloop

Fine-Tuning Is Often the Wrong First Move: Introducing Harneloop

Comments
3 min read
A LiteLLM circuit breaker that kills runaway agent runs

A LiteLLM circuit breaker that kills runaway agent runs

Comments
3 min read
I Thought Building Agent Observability Was a Detector Problem. I Was Wrong.

Real production data broke 35 detectors instantly

I Thought Building Agent Observability Was a Detector Problem. I Was Wrong.

22
Comments 19
8 min read
Our Token Counter Took 26 Seconds on a Single Prompt

Summer Bug Smash: Smash Stories 🐛🛹

Our Token Counter Took 26 Seconds on a Single Prompt

3
Comments
8 min read
I've Spent Months Grading AI Agents' Code for a Living. Here's the Pattern Nobody's Talking About

I've Spent Months Grading AI Agents' Code for a Living. Here's the Pattern Nobody's Talking About

Comments 2
4 min read
Our eval gate runs 22 minutes. The queue behind it hit three hours.

Our eval gate runs 22 minutes. The queue behind it hit three hours.

4
Comments 1
5 min read
Mem0 vs Zep vs LangChain Memory vs Letta: Picking the Right Agent Memory Architecture

Mem0 vs Zep vs LangChain Memory vs Letta: Picking the Right Agent Memory Architecture

1
Comments 4
4 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.