DEV Community

#llm

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
From a Gemma 4 Challenge Project to a Manufacturing Assistant App

From a Gemma 4 Challenge Project to a Manufacturing Assistant App

1
Comments
2 min read
AI-Cost-Explorer-Stop Guessing Your AI API Costs

AI-Cost-Explorer-Stop Guessing Your AI API Costs

Comments 3
1 min read
Grok 4.5 Shows the AI Race Is Moving From Chatbots to Agents

Grok 4.5 Shows the AI Race Is Moving From Chatbots to Agents

Comments 1
5 min read
A GitHub project claims 60-95% fewer tokens with the same answers. The number is real. The economics it implies for your agent fleet are uncomfortable.

A GitHub project claims 60-95% fewer tokens with the same answers. The number is real. The economics it implies for your agent fleet are uncomfortable.

1
Comments
14 min read
When an Actor Platform Is Too Much for an LLM Scraping Task

When an Actor Platform Is Too Much for an LLM Scraping Task

Comments
4 min read
Doubling Qwen3.6-27B on One RTX 3090: ollama llama.cpp + MTP, Lever by Lever (35.7 ~75 tok/s)

Doubling Qwen3.6-27B on One RTX 3090: ollama llama.cpp + MTP, Lever by Lever (35.7 ~75 tok/s)

Comments
7 min read
Compass v1.1.0 · we shipped a memory plugin that catches its own consumption drift

Compass v1.1.0 · we shipped a memory plugin that catches its own consumption drift

1
Comments
5 min read
OpenAI-compatible AI API gateway migration checklist

OpenAI-compatible AI API gateway migration checklist

Comments
4 min read
Design Your Own Multi-AI Coding Pipeline: A Portable Reference Architecture

Design Your Own Multi-AI Coding Pipeline: A Portable Reference Architecture

Comments
11 min read
設計你自己的 Multi-AI Coding Pipeline:一份可搬走的參考架構

設計你自己的 Multi-AI Coding Pipeline:一份可搬走的參考架構

Comments
1 min read
AI Weekly — 2026-05-29 to 2026-06-05 | The Gap Between Launch and Landing

AI Weekly — 2026-05-29 to 2026-06-05 | The Gap Between Launch and Landing

Comments
4 min read
The LLM API Failure Policy I Wish I Had Before My First Production Incident

Branching logic for RPM vs TPM rate limits

The LLM API Failure Policy I Wish I Had Before My First Production Incident

6
Comments 15
6 min read
The Data Pipeline Problems Nobody Mentions in AI Architecture Discussions

The Data Pipeline Problems Nobody Mentions in AI Architecture Discussions

1
Comments
3 min read
How to test an OpenAI-compatible AI API gateway without rewriting your app

How to test an OpenAI-compatible AI API gateway without rewriting your app

Comments
4 min read
Beyond the Model: The Hidden Cost of Multi-Model Workflows

Beyond the Model: The Hidden Cost of Multi-Model Workflows

2
Comments
1 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.