DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Testing LLMs Like Software: A Promptfoo Deep Dive for QA Engineers

Testing LLMs Like Software: A Promptfoo Deep Dive for QA Engineers

Comments 2
10 min read
Your eval suite passes. I built the tool that checks whether it checks anything.

Your eval suite passes. I built the tool that checks whether it checks anything.

3
Comments
2 min read
We almost handed out AWS keys to every team. So I built an LLM gateway instead.

We almost handed out AWS keys to every team. So I built an LLM gateway instead.

1
Comments
7 min read
Building EdgeSync-LLM: The Final Architecture for Decentralized, Offline-First Local AI 🚀

Building EdgeSync-LLM: The Final Architecture for Decentralized, Offline-First Local AI 🚀

Comments
3 min read
Harness Engineering - Part 5: Context Engineering

Harness Engineering - Part 5: Context Engineering

1
Comments
8 min read
Stop Copy-Pasting: Claude Artifacts Change the Inner Loop

Stop Copy-Pasting: Claude Artifacts Change the Inner Loop

Comments 1
3 min read
Decoupling Prompt Engineering from your Deployment Pipeline

Decoupling Prompt Engineering from your Deployment Pipeline

Comments 1
4 min read
The cheap model is only cheap for half your tasks

The cheap model is only cheap for half your tasks

1
Comments 2
4 min read
Why System Design Matters More Than Ever in the Age of LLMs

Why System Design Matters More Than Ever in the Age of LLMs

Comments
5 min read
Stop Prompting and Start Engineering: Treating LLMs as Unreliable Functions

Stop Prompting and Start Engineering: Treating LLMs as Unreliable Functions

Comments 1
3 min read
Evaluating LLMs: why 'it looks good' isn't a metric

Evaluating LLMs: why 'it looks good' isn't a metric

2
Comments 1
3 min read
AI Agents in Your Portfolio: How to Showcase Agentic Development Skills

AI Agents in Your Portfolio: How to Showcase Agentic Development Skills

Comments
3 min read
Testing Qwen3.8 Max on a Budget with OpenCode Go: 10 Tasks, ~10¢, Zero Failures

Testing Qwen3.8 Max on a Budget with OpenCode Go: 10 Tasks, ~10¢, Zero Failures

Comments
4 min read
Build a Multi-Region Canary Trap for LLM Prompt Leaks

Build a Multi-Region Canary Trap for LLM Prompt Leaks

1
Comments
4 min read
Chain-of-thought: making a model think before it answers

Chain-of-thought: making a model think before it answers

1
Comments
3 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.