DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Building an Autonomous Budget Gate: Optimizing LLM Costs with Speculative Runtime Execution

Building an Autonomous Budget Gate: Optimizing LLM Costs with Speculative Runtime Execution

Comments
3 min read
Alibaba's Open-Code-Review: A Game Changer for Code Quality?

Alibaba's Open-Code-Review: A Game Changer for Code Quality?

1
Comments 4
3 min read
Ten Layers of AI Skill Construction: A Systematic Framework from Prompts to Business Closed Loops

Ten Layers of AI Skill Construction: A Systematic Framework from Prompts to Business Closed Loops

Comments
9 min read
Hetzner Inference: First Look

Currently free with an OpenAI-compatible API

Hetzner Inference: First Look

17
Comments 4
5 min read
Fable May Not Be the Best Choice for Some Engineers

Fable May Not Be the Best Choice for Some Engineers

Comments
4 min read
22 LLM Agent Failure Modes (and the Prompts That Guard Against Them)

22 LLM Agent Failure Modes (and the Prompts That Guard Against Them)

2
Comments 3
9 min read
Your Guardrails Are a Firewall. Your Failures Are a Cascade

Your Guardrails Are a Firewall. Your Failures Are a Cascade

Comments
5 min read
AI Governance — EU AI Act Compliance, Risk Assessment, and Audit Logging

AI Governance — EU AI Act Compliance, Risk Assessment, and Audit Logging

Comments
9 min read
LLM-as-Judge Is Too Lenient. Here's a Cheap Fix: Judge Refute (Maybe) Arbitrate

LLM-as-Judge Is Too Lenient. Here's a Cheap Fix: Judge Refute (Maybe) Arbitrate

1
Comments 1
8 min read
Tiered Context Loading: Fit a Huge Agent Registry in Your Context Window

Tiered Context Loading: Fit a Huge Agent Registry in Your Context Window

Comments
7 min read
Gemma 4 E2B on a Single TPU v6e Chip: A Serving Deep Dive

Gemma 4 E2B on a Single TPU v6e Chip: A Serving Deep Dive

8
Comments 1
8 min read
The KV cache, why LLM inference is memory-bound, not compute-bound

The KV cache, why LLM inference is memory-bound, not compute-bound

Comments
4 min read
How LLM Function Calling Actually Works — From Tokens to Tool Orchestration

How LLM Function Calling Actually Works — From Tokens to Tool Orchestration

Comments
7 min read
Vision Language Models — When AI Learns to See and Talk (Part 3 of 3)

Vision Language Models — When AI Learns to See and Talk (Part 3 of 3)

Comments
13 min read
The Tool List Is the Context Window

The Tool List Is the Context Window

1
Comments 6
5 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.