DEV Community

#llm

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
The OpenAI API everyone copied isn't the one OpenAI recommends

The OpenAI API everyone copied isn't the one OpenAI recommends

2
Comments 1
7 min read
End-to-End Observability for vLLM and TGI: from DCGM to Tokens

End-to-End Observability for vLLM and TGI: from DCGM to Tokens

Comments
13 min read
AI 週報 — 2026-05-15 to 2026-05-22 | 當 IPO 傳聞撞上 27 萬人部署規模

AI 週報 — 2026-05-15 to 2026-05-22 | 當 IPO 傳聞撞上 27 萬人部署規模

Comments
3 min read
Your No-Code AI Agent Has a Memory Problem

Your No-Code AI Agent Has a Memory Problem

1
Comments
2 min read
Token economics: from model internals to agent costs

Token economics: from model internals to agent costs

Comments
5 min read
Building Production-Ready AI Applications with FastAPI and Large Language Models

Building Production-Ready AI Applications with FastAPI and Large Language Models

Comments
2 min read
Show Dev: Weavz — Governed app access for AI agents

Show Dev: Weavz — Governed app access for AI agents

1
Comments
1 min read
We Connected an LLM to a 12-Year-Old Codebase. Here's What Broke.

We Connected an LLM to a 12-Year-Old Codebase. Here's What Broke.

Comments
5 min read
Stop your AI trading agent from hallucinating technical analysis

Stop your AI trading agent from hallucinating technical analysis

Comments
2 min read
Verify the Work, Not the Report: a coding agent's success claim is just a claim

Verify the Work, Not the Report: a coding agent's success claim is just a claim

1
Comments 1
8 min read
How to Stop Evaluating LLM Outputs by Gut Feel

How to Stop Evaluating LLM Outputs by Gut Feel

Comments
4 min read
Handling Multi-Model API Outages Without Melting Production

Handling Multi-Model API Outages Without Melting Production

Comments
7 min read
Compass v1.1.0 · we shipped a memory plugin that catches its own consumption drift

Compass v1.1.0 · we shipped a memory plugin that catches its own consumption drift

1
Comments
5 min read
X's Feed Ranking Algorithm: How Grok Ranks 500M Posts in 200ms

X's Feed Ranking Algorithm: How Grok Ranks 500M Posts in 200ms

Comments
8 min read
The Request Is the Wrong Unit of Scale for LLMs on Kubernetes

The Request Is the Wrong Unit of Scale for LLMs on Kubernetes

Comments
12 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.