DEV Community

#llm

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
I streamed Mixtral 8x7B from NVMe on a $0.40/hour VM and got 3.32 tps, here's how

I streamed Mixtral 8x7B from NVMe on a $0.40/hour VM and got 3.32 tps, here's how

Comments
4 min read
How I Built a Customer Support Auto-Responder with Confidence Scoring Using pydantic-ai and FastAPI

How I Built a Customer Support Auto-Responder with Confidence Scoring Using pydantic-ai and FastAPI

Comments 2
6 min read
Agent Base Definition: Why It Is Not a Prompt

Agent Base Definition: Why It Is Not a Prompt

Comments
20 min read
Compass v1.1.0 · we shipped a memory plugin that catches its own consumption drift

Compass v1.1.0 · we shipped a memory plugin that catches its own consumption drift

1
Comments
5 min read
The Mean Is Lying to You: Benchmarks Hide the Variance That Breaks Prod

The Mean Is Lying to You: Benchmarks Hide the Variance That Breaks Prod

1
Comments
5 min read
AIchain Pool: Parallel Calls Instead of Sequential

AIchain Pool: Parallel Calls Instead of Sequential

Comments 3
5 min read
Making LLM outputs auditable: the provider abstraction pattern

Making LLM outputs auditable: the provider abstraction pattern

Comments
5 min read
Best Local Coding LLM in 2026: Qwen2.5-Coder vs DeepSeek-Coder-V2 vs Codestral

Best Local Coding LLM in 2026: Qwen2.5-Coder vs DeepSeek-Coder-V2 vs Codestral

Comments
6 min read
48,000 characters in 2,700 tokens: lets discuss how LLMs read text as images

48,000 characters in 2,700 tokens: lets discuss how LLMs read text as images

1
Comments
16 min read
API vs MCP: Understanding the Difference

API vs MCP: Understanding the Difference

1
Comments 1
3 min read
How to build cross-platform templates AI coding tools actually respect

How to build cross-platform templates AI coding tools actually respect

Comments 1
6 min read
LLM, Model, Token, Context Window

LLM, Model, Token, Context Window

Comments
6 min read
Try the Tech Radar #3 — JSON Schema LLM Prompt, Visualised

Try the Tech Radar #3 — JSON Schema LLM Prompt, Visualised

Comments
5 min read
I Got Tired of Rewriting AI API Wrappers, So I Built a Gateway

Simplified credit billing for side projects

I Got Tired of Rewriting AI API Wrappers, So I Built a Gateway

39
Comments 22
2 min read
I Built a Memory API That Beats Mem0 on LongMemEval Without Using a Single LLM Token

I Built a Memory API That Beats Mem0 on LongMemEval Without Using a Single LLM Token

1
Comments 3
4 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.