DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
The KV Cache Is the Bottleneck: A 2026 Field Guide to Attention Variants

The KV Cache Is the Bottleneck: A 2026 Field Guide to Attention Variants

1
Comments 1
4 min read
Prompt Injection Isn't Fixed by a Filter. It's Fixed by Architecture

Prompt Injection Isn't Fixed by a Filter. It's Fixed by Architecture

Comments
22 min read
Attention Sinks: Why Streaming LLMs Break When You Evict Token 0

Attention Sinks: Why Streaming LLMs Break When You Evict Token 0

Comments
6 min read
Your Pinned OpenAI Models Stop Working Next Week

Your Pinned OpenAI Models Stop Working Next Week

Comments
3 min read
Tips on Handling Files in LLMs

Tips on Handling Files in LLMs

Comments
2 min read
SDABench: A New Benchmark for Evaluating LLMs in Scientific Discovery

SDABench: A New Benchmark for Evaluating LLMs in Scientific Discovery

Comments
4 min read
Length Penalties in LLMs: Shorter Chains of Thought, Hidden Influences

Length Penalties in LLMs: Shorter Chains of Thought, Hidden Influences

Comments
3 min read
Deep Interaction: A Novel Approach to Correcting LLM Reasoning Errors

Deep Interaction: A Novel Approach to Correcting LLM Reasoning Errors

Comments
3 min read
LangChain4j and Spring AI: The Plumbing to make your Java Apps talk to LLMs

LangChain4j and Spring AI: The Plumbing to make your Java Apps talk to LLMs

Comments
8 min read
Is there a standard way to use one API key across GPT and cheaper models from other vendors in the same app?

Is there a standard way to use one API key across GPT and cheaper models from other vendors in the same app?

Comments 1
6 min read
Your cache_read_input_tokens is zero. Here is what silently did it.

Your cache_read_input_tokens is zero. Here is what silently did it.

1
Comments 2
5 min read
Getting Started with Kimi K3: API Setup, Code Examples, and First Impressions

Getting Started with Kimi K3: API Setup, Code Examples, and First Impressions

Comments
4 min read
Agentic Browser: ~98% fewer tokens than HTML for LLM web agents (Python + MCP)

Agentic Browser: ~98% fewer tokens than HTML for LLM web agents (Python + MCP)

Comments 1
1 min read
The Watermark Is the Real Boundary Between Training and Serving

The Watermark Is the Real Boundary Between Training and Serving

1
Comments
4 min read
Human-in-the-Loop Checkpoints in LangGraph

Human-in-the-Loop Checkpoints in LangGraph

Comments
3 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.