DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Gender Bias in Production LLMs: What 90 Tests Across 3 Frameworks Revealed

Gender Bias in Production LLMs: What 90 Tests Across 3 Frameworks Revealed

Comments
4 min read
From 62% to 94% RAG Accuracy: The 5 Architecture Changes That Actually Moved the Needle

From 62% to 94% RAG Accuracy: The 5 Architecture Changes That Actually Moved the Needle

1
Comments
8 min read
GPT-5 vs Claude Sonnet 4: real per-task cost and benchmark comparison for production workloads

GPT-5 vs Claude Sonnet 4: real per-task cost and benchmark comparison for production workloads

Comments 1
7 min read
Best LLM for Each Task: A Practitioner’s Reference Guide

Best LLM for Each Task: A Practitioner’s Reference Guide

Comments
13 min read
The 7 Agentic AI Design Patterns Every Developer Should Know (ReAct, Reflection, Tool Use, and More)

The 7 Agentic AI Design Patterns Every Developer Should Know (ReAct, Reflection, Tool Use, and More)

2
Comments
12 min read
Harness Engineering with Nothing but Markdown

Harness Engineering with Nothing but Markdown

1
Comments
10 min read
LLM Output Quality Metrics: How to Measure What Matters

LLM Output Quality Metrics: How to Measure What Matters

1
Comments
3 min read
GPT-5.5 Just Dropped. Here's What the Benchmarks Are Hiding.

GPT-5.5 Just Dropped. Here's What the Benchmarks Are Hiding.

1
Comments
7 min read
Epoch confirms GPT5.4 Pro solved a frontier math open problem!

Epoch confirms GPT5.4 Pro solved a frontier math open problem!

Comments
9 min read
How I accidentally built a cost tracking tool for LLMs

How I accidentally built a cost tracking tool for LLMs

Comments
2 min read
The JSON-Mode Prompt Pattern That Survives Claude Version Bumps

The JSON-Mode Prompt Pattern That Survives Claude Version Bumps

1
Comments
7 min read
Cursor Composer 2 Is a $200/Month Habit Now. Was It Worth It?

Cursor Composer 2 Is a $200/Month Habit Now. Was It Worth It?

1
Comments 1
8 min read
Why Your Shipping Speed Hasn't Changed Since You Started Using AI

Why Your Shipping Speed Hasn't Changed Since You Started Using AI

2
Comments 1
5 min read
Your RAG Eval Set Is Probably Wrong. The Test That Catches It.

Your RAG Eval Set Is Probably Wrong. The Test That Catches It.

Comments
7 min read
Stop Caching the Whole LLM Response. Cache the Embedding.

Stop Caching the Whole LLM Response. Cache the Embedding.

Comments
8 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.