DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
A Spend Cap That Stops Counting Is Already Fail-Open

A Spend Cap That Stops Counting Is Already Fail-Open

2
Comments 13
20 min read
Can You Beat an LLM? Building Humans vs. Humanity's Last Exam

Can You Beat an LLM? Building Humans vs. Humanity's Last Exam

14
Comments
4 min read
Block the Merge if the Model Isn't Ready": Shifting Local AI Evaluations Left with CI Gates

Block the Merge if the Model Isn't Ready": Shifting Local AI Evaluations Left with CI Gates

Comments
1 min read
The Harness Is Everything That Survives When the Context Forgets

The Harness Is Everything That Survives When the Context Forgets

Comments 2
9 min read
Claude AI by Anthropic: Key Features That Set This Model Apart in 2024 [EN]

Claude AI by Anthropic: Key Features That Set This Model Apart in 2024 [EN]

Comments
4 min read
I put an AI agent on a timer. Overnight it burned 136M tokens doing almost nothing.

I put an AI agent on a timer. Overnight it burned 136M tokens doing almost nothing.

Comments
4 min read
Fable 5 Banned: What Happens When Your AI Governance Lives Inside the Model

Fable 5 Banned: What Happens When Your AI Governance Lives Inside the Model

Comments 1
7 min read
Stop hand-picking an LLM per request: a practical case for auto-routing

Stop hand-picking an LLM per request: a practical case for auto-routing

Comments
3 min read
One base_url for GPT, Claude, and Gemini: cutting three SDKs down to one

One base_url for GPT, Claude, and Gemini: cutting three SDKs down to one

Comments
3 min read
OpenCode Tool Calling Internals

OpenCode Tool Calling Internals

2
Comments 3
13 min read
Build the Reply Loop: Receive, Think, Respond

Build the Reply Loop: Receive, Think, Respond

Comments
5 min read
The On-Premise LLM Lottery

The On-Premise LLM Lottery

Comments
5 min read
Stop Blaming the Model. Your Latency Budget Is Probably Broken.

Stop Blaming the Model. Your Latency Budget Is Probably Broken.

Comments
3 min read
Our few-shot examples came from the eval set. The 0.94 was fiction.

Our few-shot examples came from the eval set. The 0.94 was fiction.

1
Comments 4
13 min read
Retrieval-Augmented Self-Recall: The RAG Problem Nobody Talks About

Why similarity thresholds fail across models

Retrieval-Augmented Self-Recall: The RAG Problem Nobody Talks About

9
Comments 14
5 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.