DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Why AI Agents Fail Silently — And How to Fix It

Why AI Agents Fail Silently — And How to Fix It

1
Comments
4 min read
JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection

JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection

Comments
7 min read
One Open Source Project a Day (No.49): free-claude-code - Run Claude Code for Free with One Environment Variable

One Open Source Project a Day (No.49): free-claude-code - Run Claude Code for Free with One Environment Variable

Comments
8 min read
When Three Cheap Models Beat Claude — Through Arguing, Not Voting

When Three Cheap Models Beat Claude — Through Arguing, Not Voting

1
Comments
3 min read
How I Built an Intent Classifier to Route Messages Across Multiple LLMs

How I Built an Intent Classifier to Route Messages Across Multiple LLMs

5
Comments 3
4 min read
Best LLM for Each Task: A Practitioner’s Reference Guide

Best LLM for Each Task: A Practitioner’s Reference Guide

Comments
13 min read
Gender Bias in Production LLMs: What 90 Tests Across 3 Frameworks Revealed

Gender Bias in Production LLMs: What 90 Tests Across 3 Frameworks Revealed

Comments
4 min read
From 62% to 94% RAG Accuracy: The 5 Architecture Changes That Actually Moved the Needle

From 62% to 94% RAG Accuracy: The 5 Architecture Changes That Actually Moved the Needle

1
Comments
8 min read
The 7 Agentic AI Design Patterns Every Developer Should Know (ReAct, Reflection, Tool Use, and More)

The 7 Agentic AI Design Patterns Every Developer Should Know (ReAct, Reflection, Tool Use, and More)

2
Comments
12 min read
GPT-5 vs Claude Sonnet 4: real per-task cost and benchmark comparison for production workloads

GPT-5 vs Claude Sonnet 4: real per-task cost and benchmark comparison for production workloads

Comments 1
7 min read
LLM Output Quality Metrics: How to Measure What Matters

LLM Output Quality Metrics: How to Measure What Matters

1
Comments
3 min read
Harness Engineering with Nothing but Markdown

Harness Engineering with Nothing but Markdown

1
Comments
10 min read
GPT-5.5 Just Dropped. Here's What the Benchmarks Are Hiding.

GPT-5.5 Just Dropped. Here's What the Benchmarks Are Hiding.

1
Comments
7 min read
Epoch confirms GPT5.4 Pro solved a frontier math open problem!

Epoch confirms GPT5.4 Pro solved a frontier math open problem!

Comments
9 min read
How I accidentally built a cost tracking tool for LLMs

How I accidentally built a cost tracking tool for LLMs

Comments
2 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.