DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
OpenAI and Broadcom's Jalapeño, a Custom Inference ASIC: Inference ASIC vs GPU

OpenAI and Broadcom's Jalapeño, a Custom Inference ASIC: Inference ASIC vs GPU

5
Comments
7 min read
Tokens Are Not the Unit

Tokens Are Not the Unit

4
Comments 25
12 min read
Cheapest LLM APIs for Startups in 2026

Cheapest LLM APIs for Startups in 2026

Comments
3 min read
My eval said a perfect MCP server was broken. It was the eval that was lying.

My eval said a perfect MCP server was broken. It was the eval that was lying.

5
Comments 16
4 min read
Your AI Agent Said 'Done.' Here's How I Found Out It Actually Failed Three Hours Later.

Your AI Agent Said 'Done.' Here's How I Found Out It Actually Failed Three Hours Later.

1
Comments
5 min read
Caching in RAG Systems: What to Cache, What Not To, and Why It Matters More Than You Think

Caching in RAG Systems: What to Cache, What Not To, and Why It Matters More Than You Think

Comments 1
3 min read
I Built a Prompt Compressor That Saves 65% on LLM Costs — Here's the Story

I Built a Prompt Compressor That Saves 65% on LLM Costs — Here's the Story

Comments
2 min read
What's Next for AI?

Restricted access and global inequality

What's Next for AI?

200
Comments 257
5 min read
Testing Non-Deterministic LLM Pipelines in CI: A Contract-Based Approach

Testing Non-Deterministic LLM Pipelines in CI: A Contract-Based Approach

4
Comments 5
4 min read
How coding agents like Cursor quietly cut input costs by reusing KV states across turns — and what actually breaks the cache

How coding agents like Cursor quietly cut input costs by reusing KV states across turns — and what actually breaks the cache

5
Comments 1
4 min read
What I Learned by Deleting Rules My Agent Had Already Learned

What I Learned by Deleting Rules My Agent Had Already Learned

Comments
7 min read
Cutting our LLM bill ~80% with model routing: the actual cost math

Cutting our LLM bill ~80% with model routing: the actual cost math

Comments
3 min read
SuperCompress: Cut LLM Costs by 65% Without Losing Answers

SuperCompress: Cut LLM Costs by 65% Without Losing Answers

Comments
1 min read
Stop Prompting Your Agents. Start Leading Them.

Stop Prompting Your Agents. Start Leading Them.

Comments
9 min read
Cutting AI Costs: Batch API for Non-Urgent Workflows

Cutting AI Costs: Batch API for Non-Urgent Workflows

Comments
3 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.