DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
LLM integration with Vercel AI SDK

LLM integration with Vercel AI SDK

Comments
4 min read
Hearth: scale-to-zero LLM serving on Kubernetes — and you can hack on it without a GPU

Hearth: scale-to-zero LLM serving on Kubernetes — and you can hack on it without a GPU

2
Comments 1
3 min read
Why your quantized LLM loses its MTP heads and how to keep them

Why your quantized LLM loses its MTP heads and how to keep them

1
Comments
5 min read
Is Vibe Coding Over? The Free Lunch Is Ending, Not the Movement

Is Vibe Coding Over? The Free Lunch Is Ending, Not the Movement

Comments
4 min read
About Sharing Local Inference: A Marketplace for Renting Idle GPUs with an OpenAI-Compatible Backend

About Sharing Local Inference: A Marketplace for Renting Idle GPUs with an OpenAI-Compatible Backend

Comments
8 min read
Cut Your AI Agent Token Costs by 75% With One Skill Plugin

Cut Your AI Agent Token Costs by 75% With One Skill Plugin

Comments
2 min read
Why most LLM API usage is quietly inefficient

Why most LLM API usage is quietly inefficient

Comments
4 min read
Smarter Resource Allocation Beats Stronger Models

Smarter Resource Allocation Beats Stronger Models

Comments 2
6 min read
Qwen sky proof: compressed memory made a tiny model behave better — with the receipts

Qwen sky proof: compressed memory made a tiny model behave better — with the receipts

Comments
1 min read
Tenure — Building an AI Code Reviewer That Earns Trust Over Time

Tenure — Building an AI Code Reviewer That Earns Trust Over Time

1
Comments
4 min read
The 8B Model That Punches at 32B Weight

The 8B Model That Punches at 32B Weight

Comments
2 min read
Hermes Agent CLI cheat sheet — commands, flags, and slash shortcuts

Hermes Agent CLI cheat sheet — commands, flags, and slash shortcuts

1
Comments
8 min read
Your Agent Failed in Prod. Good Luck Reproducing It.

Recording runs over forcing determinism

Your Agent Failed in Prod. Good Luck Reproducing It.

11
Comments 13
23 min read
Running Local LLMs Without Burning Out Your GPU

Running Local LLMs Without Burning Out Your GPU

Comments
3 min read
The "Chat" API is a Token Tax: Why we must return to Stateless Completions

The "Chat" API is a Token Tax: Why we must return to Stateless Completions

Comments
2 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.