DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Your AI cited a real file. It still lied to you.

Your AI cited a real file. It still lied to you.

Comments
6 min read
An AI writes most of my publication. Its last 50 runs published nothing at all

An AI writes most of my publication. Its last 50 runs published nothing at all

5
Comments
8 min read
The Great AI Reshoring: How Silicon Valley Stacks Are Replacing the World’s Virtual Assistants

The Great AI Reshoring: How Silicon Valley Stacks Are Replacing the World’s Virtual Assistants

Comments
4 min read
GPT-6 Sol vs Opus 5.5: Price and Benchmarks

GPT-6 Sol vs Opus 5.5: Price and Benchmarks

Comments
8 min read
What Fine-Tuning an 8B Model on 250 Security Examples Actually Taught It

What Fine-Tuning an 8B Model on 250 Security Examples Actually Taught It

1
Comments 3
16 min read
LangGraph Pulled Ahead: A Code-Level Benchmark of 3 Agent Frameworks on 107 Data Engineering Tasks

LangGraph Pulled Ahead: A Code-Level Benchmark of 3 Agent Frameworks on 107 Data Engineering Tasks

Comments
5 min read
Retrieval overlap went up 13 points by promoting sentences to paragraphs

Retrieval overlap went up 13 points by promoting sentences to paragraphs

1
Comments 2
4 min read
MCP Is Dying as a Tool List. That Was Never Its Real Job.

MCP Is Dying as a Tool List. That Was Never Its Real Job.

1
Comments
5 min read
Think smaller: why specialist SLMs beat frontier models in production

Think smaller: why specialist SLMs beat frontier models in production

Comments
8 min read
Running pdlc-skills on a Real Project: Three Features, Start to Release

Running pdlc-skills on a Real Project: Three Features, Start to Release

Comments
9 min read
decider: one forward pass, typed decisions, calibrated probabilities

decider: one forward pass, typed decisions, calibrated probabilities

1
Comments 1
5 min read
Resisting Mode Gravity: Why Bigger LLMs Produce Mediocre Output

Resisting Mode Gravity: Why Bigger LLMs Produce Mediocre Output

Comments
9 min read
Designing an eval harness for prompt-injection detection: what measuring my defenses actually taught me

Designing an eval harness for prompt-injection detection: what measuring my defenses actually taught me

Comments
6 min read
Jev is now open to everyone: what a "System One" model costs, and how to start with $5 in free credit

Jev is now open to everyone: what a "System One" model costs, and how to start with $5 in free credit

Comments
1 min read
Treat streamed LLM output as an uncommitted draft

Treat streamed LLM output as an uncommitted draft

Comments
4 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.