DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Built an open-source memory layer for local LLMs — single-shot calls, auto-extracted constraints, no context degradation

Built an open-source memory layer for local LLMs — single-shot calls, auto-extracted constraints, no context degradation

Comments
1 min read
MAI-Thinking-1: Microsoft's New Reasoning Model and What It Means for Developers

MAI-Thinking-1: Microsoft's New Reasoning Model and What It Means for Developers

7
Comments 1
6 min read
Building WeaveLLM: Why .NET Deserves a Better then LangChain

Building WeaveLLM: Why .NET Deserves a Better then LangChain

Comments
9 min read
My agent swarm had a productive night. My pipeline lied about it.

My agent swarm had a productive night. My pipeline lied about it.

1
Comments 1
5 min read
How LLMs Actually Work: A Developer's Mental Model

How LLMs Actually Work: A Developer's Mental Model

1
Comments
6 min read
What Happens When You Evaluate a B2B Sales Agent on Tasks It Was Never Designed For

What Happens When You Evaluate a B2B Sales Agent on Tasks It Was Never Designed For

Comments
5 min read
Vibe Coding vs Prompt Engineering vs Context Engineering — What's the Difference?

Vibe Coding vs Prompt Engineering vs Context Engineering — What's the Difference?

Comments 1
4 min read
I tested 4 free 70B-class LLM endpoints for real production work — here's what each is actually good at

I tested 4 free 70B-class LLM endpoints for real production work — here's what each is actually good at

Comments
5 min read
Hybrid LLM Routing: Ollama + Claude API Without Quality Degradation

Hybrid LLM Routing: Ollama + Claude API Without Quality Degradation

Comments
4 min read
I tested 4 local models as memory classifiers for OpenClaw — and thinking models are a trap

I tested 4 local models as memory classifiers for OpenClaw — and thinking models are a trap

Comments
5 min read
Gemma 4 12B: Google's encoder-free multimodal AI now runs on a laptop

Gemma 4 12B: Google's encoder-free multimodal AI now runs on a laptop

1
Comments
2 min read
Helicone is now in maintenance mode. Here is how to switch to a self-hosted alternative in 5 minutes.

Helicone is now in maintenance mode. Here is how to switch to a self-hosted alternative in 5 minutes.

Comments
2 min read
KVQuant: real terminal proof for KV-cache compression

KVQuant: real terminal proof for KV-cache compression

Comments
5 min read
How to access DeepSeek and Qwen alongside OpenAI without managing separate API keys for everything

How to access DeepSeek and Qwen alongside OpenAI without managing separate API keys for everything

Comments
2 min read
Tenacious-Bench v0.1: a small B2B sales-outreach benchmark with contamination checks

Tenacious-Bench v0.1: a small B2B sales-outreach benchmark with contamination checks

Comments
2 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.