DEV Community

#llm

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
When to Move Beyond LiteLLM (And When Not To)

When to Move Beyond LiteLLM (And When Not To)

1
Comments
6 min read
Claude Code Source Analysis Series, Chapter 5: Tools Overview

Claude Code Source Analysis Series, Chapter 5: Tools Overview

1
Comments
12 min read
Tool-Response Engineering: The Frontier Beyond Prompt Engineering

Tool-Response Engineering: The Frontier Beyond Prompt Engineering

Comments
17 min read
From Theory to the Floor: What Happens When "Specificity-as-Integrity" Meets a Real Restaurant

From Theory to the Floor: What Happens When "Specificity-as-Integrity" Meets a Real Restaurant

Comments
4 min read
Most RAG failures don’t crash. They silently return bad answers. I built a repair layer for that.

Most RAG failures don’t crash. They silently return bad answers. I built a repair layer for that.

Comments
1 min read
我花了一下午测试了 NeuralBridge SDK:762KB 的 LLM 自愈方案,能用吗?

我花了一下午测试了 NeuralBridge SDK:762KB 的 LLM 自愈方案,能用吗?

Comments
1 min read
On-device LLM on iPhone: which runtime is fastest? MLX vs llama.cpp vs LiteRT-LM vs CoreML

On-device LLM on iPhone: which runtime is fastest? MLX vs llama.cpp vs LiteRT-LM vs CoreML

1
Comments 1
4 min read
Gemma 4: Frontier AI in Your Hands”

Gemma 4: Frontier AI in Your Hands”

1
Comments
2 min read
I Built a Local LLM Rig to Escape API Bills. Then I Paid OpenAI Again.

I Built a Local LLM Rig to Escape API Bills. Then I Paid OpenAI Again.

Comments 2
1 min read
BeeLlama.cpp enhances llama.cpp, Qwen 35B hits 128K context, iOS local LLMs with Ollama

BeeLlama.cpp enhances llama.cpp, Qwen 35B hits 128K context, iOS local LLMs with Ollama

Comments
3 min read
The expensive part of an AI agent failure is usually the retry loop

The expensive part of an AI agent failure is usually the retry loop

Comments 1
3 min read
AI Evals, Part 2: Error Analysis The Unglamorous Superpower Behind Good Evals

AI Evals, Part 2: Error Analysis The Unglamorous Superpower Behind Good Evals

Comments
5 min read
Deterministic reliability stack for LLM pipelines

Deterministic reliability stack for LLM pipelines

Comments
1 min read
LLM Token Counting and Cost Optimization: A Practical Guide

LLM Token Counting and Cost Optimization: A Practical Guide

1
Comments
5 min read
Generation 1 — Standalone Models (2018–2022)

Generation 1 — Standalone Models (2018–2022)

Comments
5 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.