DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Setting up LiteLLM (SDK + Proxy Gateway)

Setting up LiteLLM (SDK + Proxy Gateway)

1
Comments
2 min read
Training LLMs in the kernel — how IONA AI does embedding, RAG, and fine‑tuning without the cloud

Training LLMs in the kernel — how IONA AI does embedding, RAG, and fine‑tuning without the cloud

1
Comments
4 min read
The Day My Research Assistant Finally Got a Memory

The Day My Research Assistant Finally Got a Memory

Comments
6 min read
OK, DEGRADED, ABSTAINED, FAILED: Designing a Typed Outcome Model for LLM Calls

OK, DEGRADED, ABSTAINED, FAILED: Designing a Typed Outcome Model for LLM Calls

1
Comments 2
9 min read
3 Open-Source Repos That Each Kill a Different AI Bill(token, infra, creative)

3 Open-Source Repos That Each Kill a Different AI Bill(token, infra, creative)

Comments
4 min read
The Feynman Technique Prompt: How to Make AI Explain Anything in 4 Layers of Depth

The Feynman Technique Prompt: How to Make AI Explain Anything in 4 Layers of Depth

Comments
13 min read
Qwen3 vs DeepSeek R1: Which Open-Source Reasoning Model Should You Use in 2026?

Qwen3 vs DeepSeek R1: Which Open-Source Reasoning Model Should You Use in 2026?

Comments
8 min read
I built an RL-native memory engine for LLM agents (every memory decision becomes a training signal)

I built an RL-native memory engine for LLM agents (every memory decision becomes a training signal)

Comments
2 min read
Your Multi-Agent Pipeline Isn't Slow Because of the Model

Your Multi-Agent Pipeline Isn't Slow Because of the Model

5
Comments
3 min read
The AI judge that called a half-finished audit 'exhaustive'

The AI judge that called a half-finished audit 'exhaustive'

Comments
6 min read
MLOps for LLM: A Case Study on Dresscode

MLOps for LLM: A Case Study on Dresscode

Comments
21 min read
Why KV Cache Matters — How MQA, GQA, and MLA Make LLM Inference Faster

Why KV Cache Matters — How MQA, GQA, and MLA Make LLM Inference Faster

Comments
5 min read
Write returned success. The file was never there.

Write returned success. The file was never there.

Comments
2 min read
The 80/20 Rule of AI Code: Why Production Takes 80% of Your Time

The 80/20 Rule of AI Code: Why Production Takes 80% of Your Time

Comments
5 min read
The 50% Context Tax: Why Your AI Agent's Million-Token Window Is Burning Money

The 50% Context Tax: Why Your AI Agent's Million-Token Window Is Burning Money

Comments
4 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.