DEV Community

#llm

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
Stop Spending $500/Month on API Calls: Build Your Own LLM Pipeline

Stop Spending $500/Month on API Calls: Build Your Own LLM Pipeline

Comments
3 min read
The AI Tasks Developers Trust And the Ones They Double-Check

The AI Tasks Developers Trust And the Ones They Double-Check

Comments
11 min read
27/30 Days System Design Questions!

27/30 Days System Design Questions!

1
Comments 4
2 min read
I measured MCP vs a CLI for agent search. The MCP used 17x more tokens per call.

I measured MCP vs a CLI for agent search. The MCP used 17x more tokens per call.

11
Comments 2
6 min read
Why File-to-Markdown Conversion Is Becoming an AI Input Layer

Why File-to-Markdown Conversion Is Becoming an AI Input Layer

Comments 1
7 min read
Function-calling eval was a 2024 problem. Tool-using agents are the 2026 one.

Function-calling eval was a 2024 problem. Tool-using agents are the 2026 one.

1
Comments
5 min read
Running 35B–400B LLMs on a GPU-less Cluster to Mine 10,000 Papers — and the 4 Bugs That Almost Ruined the Data

Running 35B–400B LLMs on a GPU-less Cluster to Mine 10,000 Papers — and the 4 Bugs That Almost Ruined the Data

1
Comments
9 min read
System Prompt vs User Prompt: The Layer Under GenAI Features

System Prompt vs User Prompt: The Layer Under GenAI Features

Comments
9 min read
What makes AI API spend chargeback-safe by team/service?

What makes AI API spend chargeback-safe by team/service?

Comments
2 min read
I Compressed GPT-2 to Run on an Arduino

I Compressed GPT-2 to Run on an Arduino

Comments
1 min read
Agent Series (11): A2A Protocol — How Agents Collaborate with Each Other

Agent Series (11): A2A Protocol — How Agents Collaborate with Each Other

Comments
5 min read
Sonnet hallucinated. My agent stored it as fact.

Sonnet hallucinated. My agent stored it as fact.

3
Comments 45
3 min read
TurboQuant on a MacBook Pro, part 2: perplexity, KL divergence, and asymmetric K/V on M5 Max

TurboQuant on a MacBook Pro, part 2: perplexity, KL divergence, and asymmetric K/V on M5 Max

Comments
8 min read
llama.cpp b9455 Finally Caught vLLM: 70t/s on 2x3090 Qwen 27B UQ8

llama.cpp b9455 Finally Caught vLLM: 70t/s on 2x3090 Qwen 27B UQ8

Comments 1
3 min read
Why I'm Building a Local-First AI Coding Workspace (And How Behavioral Routing Makes It Work)

Why I'm Building a Local-First AI Coding Workspace (And How Behavioral Routing Makes It Work)

Comments
6 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.