DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
The hard part of agent memory isn't remembering — it's forgetting

The hard part of agent memory isn't remembering — it's forgetting

1
Comments
4 min read
The LLM Kept Saying “Fixed.” For Three Months, It Wasn’t.

The LLM Kept Saying “Fixed.” For Three Months, It Wasn’t.

Comments
7 min read
Inference Arbitrage: How I Route 200+ Daily LLM Calls Across Five Models

Inference Arbitrage: How I Route 200+ Daily LLM Calls Across Five Models

Comments
10 min read
Three Months of Speed-Up Experiments on a 3090 Ti: Autoregressive DFlash MTP for Qwen3.6-27B

Three Months of Speed-Up Experiments on a 3090 Ti: Autoregressive DFlash MTP for Qwen3.6-27B

Comments
18 min read
Building llama.cpp from source on a Dell Precision T5820 with an RTX 3090 Ti (after seven power cycles)

Building llama.cpp from source on a Dell Precision T5820 with an RTX 3090 Ti (after seven power cycles)

Comments
16 min read
How I Track Claude, Codex, and Gemini Quotas from One Script

How I Track Claude, Codex, and Gemini Quotas from One Script

Comments
6 min read
Designing a Multi-Agent AI System for Content Analysis and Recommendations

Designing a Multi-Agent AI System for Content Analysis and Recommendations

Comments
7 min read
Claude Mythos vs Opus 4.8: 90x More Firefox Exploits — But Stay on Opus Anyway

Claude Mythos vs Opus 4.8: 90x More Firefox Exploits — But Stay on Opus Anyway

5
Comments
6 min read
I Cut My LLM API Bill by 73% — Here's the Exact Optimization Playbook

I Cut My LLM API Bill by 73% — Here's the Exact Optimization Playbook

Comments
5 min read
Why Your Reranker Isn't Helping Your RAG Pipeline (And How to Prove It)

Why Your Reranker Isn't Helping Your RAG Pipeline (And How to Prove It)

1
Comments 4
4 min read
What Production ML Systems Taught Me About AI Hallucinations

What Production ML Systems Taught Me About AI Hallucinations

Comments
4 min read
AI Red-Teaming Techniques: A Practical Starting Point for Security Teams

AI Red-Teaming Techniques: A Practical Starting Point for Security Teams

Comments 1
4 min read
Local Inference Boost: Qwen 3.6 Benchmarks, KV Cache Quantization, & Ollama UI

Local Inference Boost: Qwen 3.6 Benchmarks, KV Cache Quantization, & Ollama UI

Comments
3 min read
Kimi K2.6 Beats Frontier Models in Coding Benchmarks

Kimi K2.6 Beats Frontier Models in Coding Benchmarks

Comments
6 min read
From Burnout to Building: One Indie Dev's Story Behind Mozart

From Burnout to Building: One Indie Dev's Story Behind Mozart

Comments
5 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.