DEV Community

#llm

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
Routing diffusion inference traffic across three providers

Routing diffusion inference traffic across three providers

Comments
4 min read
ToolRouter: Switch AI Coding Tools Freely Without Losing Context

ToolRouter: Switch AI Coding Tools Freely Without Losing Context

2
Comments
6 min read
Beyond the Stateless Prompt: Building an Auditable Product Intelligence Pipeline with Cascadeflow and Hindsight

Beyond the Stateless Prompt: Building an Auditable Product Intelligence Pipeline with Cascadeflow and Hindsight

Comments
5 min read
Putting an LLM Gateway in Front of Our Build Agents

Putting an LLM Gateway in Front of Our Build Agents

Comments
4 min read
You Probably Don't Need 8-Bit Quantization

You Probably Don't Need 8-Bit Quantization

Comments
2 min read
Five ways your AI coding agent wastes tokens (and how to fix each one)

Five ways your AI coding agent wastes tokens (and how to fix each one)

2
Comments 1
6 min read
OpenAI-Compatible API Gateway: Base URL, API Keys, and Model Routing for Dify, Cursor, and Node.js

OpenAI-Compatible API Gateway: Base URL, API Keys, and Model Routing for Dify, Cursor, and Node.js

1
Comments
7 min read
The README Was a Protocol. The Entrypoint Was Still Optional.

The README Was a Protocol. The Entrypoint Was Still Optional.

Comments
8 min read
The OpenAI API everyone copied isn't the one OpenAI recommends

The OpenAI API everyone copied isn't the one OpenAI recommends

2
Comments 1
7 min read
End-to-End Observability for vLLM and TGI: from DCGM to Tokens

End-to-End Observability for vLLM and TGI: from DCGM to Tokens

Comments
13 min read
AI 週報 — 2026-05-15 to 2026-05-22 | 當 IPO 傳聞撞上 27 萬人部署規模

AI 週報 — 2026-05-15 to 2026-05-22 | 當 IPO 傳聞撞上 27 萬人部署規模

Comments
3 min read
Your No-Code AI Agent Has a Memory Problem

Your No-Code AI Agent Has a Memory Problem

1
Comments
2 min read
Token economics: from model internals to agent costs

Token economics: from model internals to agent costs

Comments
5 min read
Why RAG Isn't Enough: Building RationaleVault for Cognitive Continuity

Why RAG Isn't Enough: Building RationaleVault for Cognitive Continuity

Comments 1
4 min read
Building Production-Ready AI Applications with FastAPI and Large Language Models

Building Production-Ready AI Applications with FastAPI and Large Language Models

Comments
2 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.