DEV Community

#llm

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
Cut LLM prompt tokens on structured data — losslessly

Cut LLM prompt tokens on structured data — losslessly

1
Comments 5
3 min read
llama.cpp Native Tools, Qwen GGUF Models, and Local Multimodal Audio Tools

llama.cpp Native Tools, Qwen GGUF Models, and Local Multimodal Audio Tools

Comments
3 min read
로컬 LLM 셋업 가이드 (v21)

로컬 LLM 셋업 가이드 (v21)

Comments
2 min read
How Claude Code Achieves a 92% Cache Hit Rate: A Deep Dive Into Prompt Caching for AI Agents

How Claude Code Achieves a 92% Cache Hit Rate: A Deep Dive Into Prompt Caching for AI Agents

Comments
8 min read
DeepSeek's DSpark Brings Speculative Decoding Back Into the Spotlight — Here's What Developers Need to Know

DeepSeek's DSpark Brings Speculative Decoding Back Into the Spotlight — Here's What Developers Need to Know

1
Comments 1
4 min read
터미널 AI 에이전트 구축 (v20)

터미널 AI 에이전트 구축 (v20)

Comments
3 min read
Agentic AI Search

Agentic AI Search

Comments
11 min read
Persistent memory for Ollama, in about five minutes

Persistent memory for Ollama, in about five minutes

3
Comments 4
5 min read
prompt-shield: a tiny, zero-dep prompt-injection detector you can drop in front of any agent

prompt-shield: a tiny, zero-dep prompt-injection detector you can drop in front of any agent

Comments
5 min read
Calibrated LLM-as-judge: how I made my LLM give honest 4/10 scores instead of always-an-8

Calibrated LLM-as-judge: how I made my LLM give honest 4/10 scores instead of always-an-8

Comments
6 min read
로컬 LLM 셋업 가이드 (v18)

로컬 LLM 셋업 가이드 (v18)

Comments
3 min read
Building a Markdown-to-JSON Pipeline with Structured LLM Output

Building a Markdown-to-JSON Pipeline with Structured LLM Output

2
Comments
4 min read
When the Free Executor Cost More: 40 Trials on Opus + Local Qwen Ended Up the Most Expensive Cloud Arm

When the Free Executor Cost More: 40 Trials on Opus + Local Qwen Ended Up the Most Expensive Cloud Arm

2
Comments 1
8 min read
Building a RAG System from Scratch — AI Agents: Memory, Planning, and Multi-Step Reasoning

Building a RAG System from Scratch — AI Agents: Memory, Planning, and Multi-Step Reasoning

Comments 1
7 min read
Implementing Token Budgets to Prevent AI Cost Overruns

Implementing Token Budgets to Prevent AI Cost Overruns

Comments 1
3 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.