DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
New `llama.cpp` Updates, AI Agents for Any LLM, and Quantized Vector Index for Local Inference

New `llama.cpp` Updates, AI Agents for Any LLM, and Quantized Vector Index for Local Inference

Comments
3 min read
I'm 13, Building a CLI Tool for LLM Cost Tracking, and Shipping It in 10 Days

Privacy-first local logging with no account

I'm 13, Building a CLI Tool for LLM Cost Tracking, and Shipping It in 10 Days

8
Comments 9
3 min read
I wired 908 creator dossiers into my Substack commenter. Here is what changed.

I wired 908 creator dossiers into my Substack commenter. Here is what changed.

Comments
3 min read
The 20% of your AI agent's tool schemas that's pure cruft (and the one-liner to strip it)

The 20% of your AI agent's tool schemas that's pure cruft (and the one-liner to strip it)

Comments
2 min read
DeepSeek V4 Pro vs MiMo V2.5 Pro - Debugging Benchmark

DeepSeek V4 Pro vs MiMo V2.5 Pro - Debugging Benchmark

Comments 1
6 min read
We stopped Googling and started Prompting

We stopped Googling and started Prompting

Comments
4 min read
Prompt injection and LLM security for SaaS

Prompt injection and LLM security for SaaS

1
Comments
10 min read
Compass v1.1.0 · we shipped a memory plugin that catches its own consumption drift

Compass v1.1.0 · we shipped a memory plugin that catches its own consumption drift

Comments
5 min read
Running Local LLMs Without Burning Out Your GPU

Running Local LLMs Without Burning Out Your GPU

Comments
3 min read
From Code Completion to Autonomous Reasoning: What the Oceanus Leak Tells Us About the Future of AI Software Engineering

From Code Completion to Autonomous Reasoning: What the Oceanus Leak Tells Us About the Future of AI Software Engineering

Comments
7 min read
How to Handle LLM API Errors & Rate Limits in Node.js

How to Handle LLM API Errors & Rate Limits in Node.js

Comments
4 min read
I measured the token cost of 13 real AI agents (GitHub's MCP server alone is 3,546 tokens/turn)

I measured the token cost of 13 real AI agents (GitHub's MCP server alone is 3,546 tokens/turn)

Comments 1
2 min read
AIchain Reasoning: One Parameter for Every Provider

AIchain Reasoning: One Parameter for Every Provider

Comments
5 min read
MarginGate: Margin-Gated Verification for Batch-Invariant Decoding

MarginGate: Margin-Gated Verification for Batch-Invariant Decoding

Comments
5 min read
redb.Route.Llm 3.1.1 — per-message audit fields for LLM compliance / replay

redb.Route.Llm 3.1.1 — per-message audit fields for LLM compliance / replay

Comments
1 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.