DEV Community

#llm

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
The Circuit Breaker Pattern for AI Agents

The Circuit Breaker Pattern for AI Agents

7
Comments 2
9 min read
How I reduced LLM context cost by 35% without changing code (Token Firewall)

How I reduced LLM context cost by 35% without changing code (Token Firewall)

1
Comments 1
2 min read
Tokenization in AI: What Is It, Why Is It Used, and How Does It Work?

Tokenization in AI: What Is It, Why Is It Used, and How Does It Work?

Comments
6 min read
LLM中如果一个问题容易验证 那么AI就容易学会解决!说说这个特性与P与NP问题的关联性

LLM中如果一个问题容易验证 那么AI就容易学会解决!说说这个特性与P与NP问题的关联性

Comments
1 min read
The Open-Weight Inflection Point: Kimi K3, Claude Opus 5, and Microsoft MAI Signal a Market Shift

The Open-Weight Inflection Point: Kimi K3, Claude Opus 5, and Microsoft MAI Signal a Market Shift

Comments
3 min read
I built a Laravel package for LLM workflows with hard token and cost limits

I built a Laravel package for LLM workflows with hard token and cost limits

Comments
2 min read
Micro-compaction: amortizing context compression in agent loops

Micro-compaction: amortizing context compression in agent loops

1
Comments 3
5 min read
Kimi K3 is the largest open-weight model ever released — and you probably still can't run it

Kimi K3 is the largest open-weight model ever released — and you probably still can't run it

7
Comments
2 min read
Benchmarking GPT-4o, Claude 3.5 Sonnet, and Llama 3 for Automated Code Auditing & Vulnerability Detection

Benchmarking GPT-4o, Claude 3.5 Sonnet, and Llama 3 for Automated Code Auditing & Vulnerability Detection

Comments
2 min read
Same DeepSeek V4 Flash, Different Agent: Why the Runtime Changes the Result

Same DeepSeek V4 Flash, Different Agent: Why the Runtime Changes the Result

Comments
2 min read
How I Use DeepSeek V4 Flash: Reserve the Strongest Model for Uncertainty

How I Use DeepSeek V4 Flash: Reserve the Strongest Model for Uncertainty

Comments
2 min read
Building a podcast summarizer in 20 lines of Python

Building a podcast summarizer in 20 lines of Python

Comments
3 min read
Building Sluice: QoS-Aware Capacity Governance for Self-Hosted LLM Inference

Building Sluice: QoS-Aware Capacity Governance for Self-Hosted LLM Inference

1
Comments 1
19 min read
Stop Your AI Coding CLI From Wasting Tokens on "Hi" and "Thanks"

Stop Your AI Coding CLI From Wasting Tokens on "Hi" and "Thanks"

5
Comments 5
6 min read
Ingest-Time Compilation Takes On Query-Time RAG, and Agentic Retrieval Meets Its Limits

Ingest-Time Compilation Takes On Query-Time RAG, and Agentic Retrieval Meets Its Limits

1
Comments
6 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.