DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
llama.cpp Checkpoint Fix, NuExtract3 VLM, & Qwen3.6 Local Inference Benchmarks

llama.cpp Checkpoint Fix, NuExtract3 VLM, & Qwen3.6 Local Inference Benchmarks

Comments
3 min read
Your GitHub Actions Logs Are Leaking LLM Keys and Your SIEM Isn't Catching It

Your GitHub Actions Logs Are Leaking LLM Keys and Your SIEM Isn't Catching It

Comments
3 min read
The Human in the Loop Doesn't Scale. I Kept Him Anyway.

The Human in the Loop Doesn't Scale. I Kept Him Anyway.

1
Comments 2
6 min read
Auto-labelling 1.2M robotics frames with VLMs: a failover story

Auto-labelling 1.2M robotics frames with VLMs: a failover story

Comments
4 min read
My AI agent got dumber mid-session. I measured the context window before blaming MCP.

History bloat vs minimal tool overhead

My AI agent got dumber mid-session. I measured the context window before blaming MCP.

21
Comments 85
5 min read
My RAG Benchmark is lying to me

My RAG Benchmark is lying to me

1
Comments 1
5 min read
Getting structured JSON out of five incompatible LLM APIs — and degrading when they ignore you

The parser as the real system contract

Getting structured JSON out of five incompatible LLM APIs — and degrading when they ignore you

7
Comments 13
5 min read
We Audited Our Agent Tool-Call Traces. Half Our Eval Data Was Garbage.

We Audited Our Agent Tool-Call Traces. Half Our Eval Data Was Garbage.

Comments
4 min read
KV Cache Is Eating Your VRAM — Here's How to Estimate It Before You Run Out

KV Cache Is Eating Your VRAM — Here's How to Estimate It Before You Run Out

Comments
6 min read
Gemma 4: Google's Lightweight Powerhouse — Run AI on Hardware You Already Own

Gemma 4: Google's Lightweight Powerhouse — Run AI on Hardware You Already Own

Comments
3 min read
GLM-4: The Chinese-English Bilingual Workhorse You Didn't Know You Needed

GLM-4: The Chinese-English Bilingual Workhorse You Didn't Know You Needed

Comments
3 min read
Trust Isn't a Scalar: Typed Provenance for Agent Chains

Co-authored by the article's comment section

Trust Isn't a Scalar: Typed Provenance for Agent Chains

14
Comments 27
9 min read
Preventing GPT hallucination in automated content pipelines: how I structure Make.com flows with data injection

Preventing GPT hallucination in automated content pipelines: how I structure Make.com flows with data injection

Comments
8 min read
I ran Claude Code on a local LLM for 4 hours — 7M tokens, $0 (would have cost $94)

I ran Claude Code on a local LLM for 4 hours — 7M tokens, $0 (would have cost $94)

Comments
2 min read
Cost accounting for diffusion image generation at $0.0008 per render

Cost accounting for diffusion image generation at $0.0008 per render

Comments
4 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.