DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
cachebench: stop finding out about prompt-cache regressions from the invoice

cachebench: stop finding out about prompt-cache regressions from the invoice

Comments
4 min read
I needed a stable cache key for LLM requests. The hard part was the input list order.

I needed a stable cache key for LLM requests. The hard part was the input list order.

Comments
4 min read
My LLM provider went down for 11 minutes. My code spent 4 of them in connect timeouts.

My LLM provider went down for 11 minutes. My code spent 4 of them in connect timeouts.

Comments
4 min read
Compass v1.1.0 · we shipped a memory plugin that catches its own consumption drift

Compass v1.1.0 · we shipped a memory plugin that catches its own consumption drift

1
Comments
5 min read
We Let Sci-Fi Authors Code AI For Us

We Let Sci-Fi Authors Code AI For Us

2
Comments
4 min read
.klickd v4.0.0 — Portable AI memory with constraints, strict schemas, and test vectors

.klickd v4.0.0 — Portable AI memory with constraints, strict schemas, and test vectors

Comments
6 min read
What 'Bring Your Own Model' (BYOK) Actually Means When You Adopt AI at Work

What 'Bring Your Own Model' (BYOK) Actually Means When You Adopt AI at Work

Comments
4 min read
Trust Isn't a Scalar: Typed Provenance for Agent Chains

Co-authored by the article's comment section

Trust Isn't a Scalar: Typed Provenance for Agent Chains

14
Comments 27
9 min read
Your GitHub Actions Logs Are Leaking LLM Keys and Your SIEM Isn't Catching It

Your GitHub Actions Logs Are Leaking LLM Keys and Your SIEM Isn't Catching It

Comments
3 min read
Lossless, But Not Free: The Lossless, But Not Free — When Speculative Decoding Actually Pays Off (and When It Doesn't)

Lossless, But Not Free: The Lossless, But Not Free — When Speculative Decoding Actually Pays Off (and When It Doesn't)

2
Comments 4
6 min read
llama.cpp Checkpoint Fix, NuExtract3 VLM, & Qwen3.6 Local Inference Benchmarks

llama.cpp Checkpoint Fix, NuExtract3 VLM, & Qwen3.6 Local Inference Benchmarks

Comments
3 min read
How to Build a Capability-First Model Portfolio in TypeScript

How to Build a Capability-First Model Portfolio in TypeScript

1
Comments
1 min read
Auto-labelling 1.2M robotics frames with VLMs: a failover story

Auto-labelling 1.2M robotics frames with VLMs: a failover story

Comments
4 min read
My AI agent got dumber mid-session. I measured the context window before blaming MCP.

History bloat vs minimal tool overhead

My AI agent got dumber mid-session. I measured the context window before blaming MCP.

21
Comments 85
5 min read
My RAG Benchmark is lying to me

My RAG Benchmark is lying to me

1
Comments 1
5 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.