DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Portable AI on a USB Stick

Portable AI on a USB Stick

Comments
1 min read
How llm-d Prefix-Cache Routing Made Qwen 7B on EKS 2.3x Faster

How llm-d Prefix-Cache Routing Made Qwen 7B on EKS 2.3x Faster

Comments
7 min read
Why Kimi K3 Still Can't Do What Einstein Did

RAG surfaces echoes, but misses paradigm shifts

Why Kimi K3 Still Can't Do What Einstein Did

35
Comments 22
4 min read
Auto-Compact Ate My Constraints: A Post-Mortem

Auto-Compact Ate My Constraints: A Post-Mortem

2
Comments 2
3 min read
How I Built a No-Execution LLM Eval Judge (75% Accuracy, No API Calls)

How I Built a No-Execution LLM Eval Judge (75% Accuracy, No API Calls)

1
Comments
2 min read
How Do AI Agents Remember Things Between Conversations?

How Do AI Agents Remember Things Between Conversations?

Comments
4 min read
AI Harnesses Are Just Middleware, and Middleware Trust Bugs Are Older Than Your Career

AI Harnesses Are Just Middleware, and Middleware Trust Bugs Are Older Than Your Career

1
Comments 2
3 min read
Bringing Notes, WeChat Reading, and Zhihu into Obsidian: My LLM-Wiki Knowledge Hub

Bringing Notes, WeChat Reading, and Zhihu into Obsidian: My LLM-Wiki Knowledge Hub

1
Comments
7 min read
How LLMs Decide Which Ecommerce Brands to Recommend, and What Shopify Stores Need to Do About It

How LLMs Decide Which Ecommerce Brands to Recommend, and What Shopify Stores Need to Do About It

Comments
4 min read
Calling a Model Is Easy. Running an AI System Is Not

Calling a Model Is Easy. Running an AI System Is Not

Comments
7 min read
Your Scaffold Will Be Gamed

Your Scaffold Will Be Gamed

1
Comments
6 min read
What Is Agentic AI? And Why Oversight Has to Change

What Is Agentic AI? And Why Oversight Has to Change

1
Comments 1
8 min read
LLMs on Consumer Hardware — Part 3: A Multi-Agent Setup with Hermes — Delegation and Early Failure Modes

LLMs on Consumer Hardware — Part 3: A Multi-Agent Setup with Hermes — Delegation and Early Failure Modes

Comments
5 min read
Your AI Agent Folds When You Push Back: Measured Sycophancy and a Challenge-Triggered Verification Gate

Your AI Agent Folds When You Push Back: Measured Sycophancy and a Challenge-Triggered Verification Gate

Comments 25
8 min read
OpenAI and Broadcom's Jalapeño, a Custom Inference ASIC: Inference ASIC vs GPU

OpenAI and Broadcom's Jalapeño, a Custom Inference ASIC: Inference ASIC vs GPU

5
Comments
7 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.