DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Hype Cycles Don't Ship My Code: How I Gate New LLMs with a Self-Written Eval Deck

Hype Cycles Don't Ship My Code: How I Gate New LLMs with a Self-Written Eval Deck

Comments
5 min read
Your AI Agent Has Access to Your Database. What Could Go Wrong?

Your AI Agent Has Access to Your Database. What Could Go Wrong?

Comments
5 min read
Your Feed Says the New Model Is Great. Mine Says Prove It in Under an Hour

Your Feed Says the New Model Is Great. Mine Says Prove It in Under an Hour

Comments
4 min read
Context Windows: Why Too Much Text Breaks AI in Production

Context Windows: Why Too Much Text Breaks AI in Production

Comments
6 min read
Pick Your LLM With a Scoreboard, Not a Hunch: A Two-File Eval That Costs Nothing to Re-Run

Pick Your LLM With a Scoreboard, Not a Hunch: A Two-File Eval That Costs Nothing to Re-Run

Comments
6 min read
How we cut our LLM bill 40% with multi-provider routing

How we cut our LLM bill 40% with multi-provider routing

Comments
2 min read
Stop Picking LLMs by Vibes: A Reproducible Evaluation Harness You Can Run for Free

Stop Picking LLMs by Vibes: A Reproducible Evaluation Harness You Can Run for Free

Comments
5 min read
'PleaseFix': When Your AI Browser Trusts a Web Page More Than It Trusts You

'PleaseFix': When Your AI Browser Trusts a Web Page More Than It Trusts You

1
Comments
5 min read
Stop Benchmarking Coding Models on Strangers' Bugs: A Reproducible Harness for Your Own Repo

Stop Benchmarking Coding Models on Strangers' Bugs: A Reproducible Harness for Your Own Repo

Comments
5 min read
I Built a Personal Regression Suite for LLMs — Here's the Design, Not Just the Code

I Built a Personal Regression Suite for LLMs — Here's the Design, Not Just the Code

Comments
7 min read
Meta Muse Glimmer: The New 30B Open Weights Coding Model!

Meta Muse Glimmer: The New 30B Open Weights Coding Model!

Comments
5 min read
From Manual Commands to Intent-Driven Maintenance: The Agentic Workflow

From Manual Commands to Intent-Driven Maintenance: The Agentic Workflow

5
Comments
4 min read
How to test your LLM app for prompt injection: promptfoo vs garak vs Giskard vs PyRIT vs sentinel-scan-cli

How to test your LLM app for prompt injection: promptfoo vs garak vs Giskard vs PyRIT vs sentinel-scan-cli

1
Comments
5 min read
You cannot predict what an LLM call will cost before you make it

You cannot predict what an LLM call will cost before you make it

Comments
4 min read
Stop Letting Claude Code Hallucinate Your Modernization: Write OpenRewrite LST Recipes Instead

Stop Letting Claude Code Hallucinate Your Modernization: Write OpenRewrite LST Recipes Instead

Comments
2 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.