DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
What 'Bring Your Own Model' (BYOK) Actually Means When You Adopt AI at Work

What 'Bring Your Own Model' (BYOK) Actually Means When You Adopt AI at Work

Comments
4 min read
.klickd v4.0.0 — Portable AI memory with constraints, strict schemas, and test vectors

.klickd v4.0.0 — Portable AI memory with constraints, strict schemas, and test vectors

Comments
6 min read
llama.cpp Checkpoint Fix, NuExtract3 VLM, & Qwen3.6 Local Inference Benchmarks

llama.cpp Checkpoint Fix, NuExtract3 VLM, & Qwen3.6 Local Inference Benchmarks

Comments
3 min read
Your GitHub Actions Logs Are Leaking LLM Keys and Your SIEM Isn't Catching It

Your GitHub Actions Logs Are Leaking LLM Keys and Your SIEM Isn't Catching It

Comments
3 min read
Lossless, But Not Free: The Lossless, But Not Free — When Speculative Decoding Actually Pays Off (and When It Doesn't)

Lossless, But Not Free: The Lossless, But Not Free — When Speculative Decoding Actually Pays Off (and When It Doesn't)

2
Comments 4
6 min read
The Human in the Loop Doesn't Scale. I Kept Him Anyway.

The Human in the Loop Doesn't Scale. I Kept Him Anyway.

1
Comments 2
6 min read
Auto-labelling 1.2M robotics frames with VLMs: a failover story

Auto-labelling 1.2M robotics frames with VLMs: a failover story

Comments
4 min read
How to Build a Capability-First Model Portfolio in TypeScript

How to Build a Capability-First Model Portfolio in TypeScript

1
Comments
1 min read
My RAG Benchmark is lying to me

My RAG Benchmark is lying to me

1
Comments 1
5 min read
My AI agent got dumber mid-session. I measured the context window before blaming MCP.

History bloat vs minimal tool overhead

My AI agent got dumber mid-session. I measured the context window before blaming MCP.

21
Comments 85
5 min read
Getting structured JSON out of five incompatible LLM APIs — and degrading when they ignore you

The parser as the real system contract

Getting structured JSON out of five incompatible LLM APIs — and degrading when they ignore you

7
Comments 13
5 min read
KV Cache Is Eating Your VRAM — Here's How to Estimate It Before You Run Out

KV Cache Is Eating Your VRAM — Here's How to Estimate It Before You Run Out

Comments
6 min read
We Audited Our Agent Tool-Call Traces. Half Our Eval Data Was Garbage.

We Audited Our Agent Tool-Call Traces. Half Our Eval Data Was Garbage.

Comments
4 min read
Gemma 4: Google's Lightweight Powerhouse — Run AI on Hardware You Already Own

Gemma 4: Google's Lightweight Powerhouse — Run AI on Hardware You Already Own

Comments
3 min read
GLM-4: The Chinese-English Bilingual Workhorse You Didn't Know You Needed

GLM-4: The Chinese-English Bilingual Workhorse You Didn't Know You Needed

Comments
3 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.