DEV Community

#llm

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
Why the E8 lattice is the perfect quantizer for KV caches

Why the E8 lattice is the perfect quantizer for KV caches

1
Comments
2 min read
What Is Context Engineering and How to Apply It in Real Systems

What Is Context Engineering and How to Apply It in Real Systems

4
Comments
6 min read
How I Built a PII Tokenization Middleware to Keep Sensitive Data Out of LLM APIs

How I Built a PII Tokenization Middleware to Keep Sensitive Data Out of LLM APIs

9
Comments 4
5 min read
The 12 approaches I tested before finding one that works

The 12 approaches I tested before finding one that works

Comments
5 min read
I Benchmarked 5 File Editing Strategies for AI Coding Agents. Here's What Actually Works.

I Benchmarked 5 File Editing Strategies for AI Coding Agents. Here's What Actually Works.

2
Comments 3
2 min read
How I Built Persistent Memory for Claude Code

Uses all six Claude Code hooks

How I Built Persistent Memory for Claude Code

10
Comments 32
9 min read
RAG in the Wild: What I Learned After Two Weeks of Chunking Experiments

RAG in the Wild: What I Learned After Two Weeks of Chunking Experiments

Comments 2
7 min read
I benchmarked identity drift across 5 AI agent memory architectures — here's what I found

I benchmarked identity drift across 5 AI agent memory architectures — here's what I found

Comments
3 min read
Running 1M-token context on a single GPU (the math)

Running 1M-token context on a single GPU (the math)

Comments
2 min read
I Read a Paper That Genuinely Made Me Stop and Think — AI is Now Jailbreaking Other AI

I Read a Paper That Genuinely Made Me Stop and Think — AI is Now Jailbreaking Other AI

Comments
3 min read
One line of Python to extend your LLM's context window 10x

One line of Python to extend your LLM's context window 10x

Comments
1 min read
Build Your Own AI-Powered Knowledge Base with LLMs and Obsidian

Build Your Own AI-Powered Knowledge Base with LLMs and Obsidian

4
Comments
6 min read
KV cache memory calculator: how much does your LLM actually use?

KV cache memory calculator: how much does your LLM actually use?

Comments
3 min read
How Much GPU Memory Does NexusQuant Actually Save?

How Much GPU Memory Does NexusQuant Actually Save?

Comments
4 min read
The Math Behind E8 Lattice Quantization (with Code)

The Math Behind E8 Lattice Quantization (with Code)

Comments
6 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.