DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Qwen3.6-27B + vLLM + Hermes on 24GB VRAM: May 2026 Recipe

Qwen3.6-27B + vLLM + Hermes on 24GB VRAM: May 2026 Recipe

1
Comments
4 min read
Stop reinventing 'ask GPT-4 and Claude and a regex, then count the votes'

Stop reinventing 'ask GPT-4 and Claude and a regex, then count the votes'

Comments
4 min read
The Agent Harness Is the Real Product. The Model Is Just the Engine.

The Agent Harness Is the Real Product. The Model Is Just the Engine.

Comments
10 min read
Your Agent Doesn't Run Out of Context. It Degrades at 79%

Your Agent Doesn't Run Out of Context. It Degrades at 79%

1
Comments 2
8 min read
Your agent's memory should compute confidence, not store it

Your agent's memory should compute confidence, not store it

1
Comments 9
4 min read
I Replaced My Old AI Agent Stack with Hermes — and It Finally Feels Like Infrastructure

I Replaced My Old AI Agent Stack with Hermes — and It Finally Feels Like Infrastructure

Comments 2
3 min read
RLHF in 2026: when to pick PPO, DPO, or verifier-based RL

RLHF in 2026: when to pick PPO, DPO, or verifier-based RL

Comments
7 min read
Stop hallucinating: a developer API for grounding LLM responses with signed, sourced claims

Stop hallucinating: a developer API for grounding LLM responses with signed, sourced claims

Comments
4 min read
Why Does AI Have Limits? Understanding What Today's Models Can't Do

Why Does AI Have Limits? Understanding What Today's Models Can't Do

3
Comments
4 min read
I am new here

I am new here

Comments
1 min read
LLM providers are retiring models faster than you can migrate

LLM providers are retiring models faster than you can migrate

Comments
2 min read
Inside Systems 01: Your Verification Process Did Not Break. It Was Replaced.

Inside Systems 01: Your Verification Process Did Not Break. It Was Replaced.

Comments
8 min read
Your LLM provider will deprecate your model. xAI just gave 9 days' notice.

Your LLM provider will deprecate your model. xAI just gave 9 days' notice.

Comments
2 min read
Running Local GGUF Models with Ollama (GPU Enabled)

Running Local GGUF Models with Ollama (GPU Enabled)

Comments
3 min read
Context Engineering: Building More Reliable LLM Systems in Production

Context Engineering: Building More Reliable LLM Systems in Production

Comments
3 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.