DEV Community

#llm

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
TitanCore Core-1 – Trillion-parameter LLM training infra in C++/CUDA with ZeRO-3

TitanCore Core-1 – Trillion-parameter LLM training infra in C++/CUDA with ZeRO-3

Comments
1 min read
Building and Running Llama.cpp on an Air-Gapped Mac

Building and Running Llama.cpp on an Air-Gapped Mac

Comments
3 min read
AIMO: AI Mention Optimization — The Discipline of Being Recommended by AI Assistants

AIMO: AI Mention Optimization — The Discipline of Being Recommended by AI Assistants

Comments
6 min read
从 pip install 到生产部署:AI 自愈 Agent 10 分钟上线指南

从 pip install 到生产部署:AI 自愈 Agent 10 分钟上线指南

Comments
2 min read
Multi-Agent Kill Switch: Why Stopping the Orchestrator Doesn't Stop the Swarm

Multi-Agent Kill Switch: Why Stopping the Orchestrator Doesn't Stop the Swarm

1
Comments 1
11 min read
GroundedQL: a semantic compiler for natural-language Postgres analytics

GroundedQL: a semantic compiler for natural-language Postgres analytics

Comments
2 min read
llama.cpp Optimizations & New Qwopus3.5-9B GGUF Model Boost Local AI Performance

llama.cpp Optimizations & New Qwopus3.5-9B GGUF Model Boost Local AI Performance

Comments
3 min read
Why I used three different critic roles instead of one (and what the eval taught me)

Why I used three different critic roles instead of one (and what the eval taught me)

Comments 2
6 min read
Fitting LLM Reply Suggestions Into Every Provider's Prompt Cache — Without Structured Output

Fitting LLM Reply Suggestions Into Every Provider's Prompt Cache — Without Structured Output

Comments 1
4 min read
High-Value If, Low-Value Foreach: Why Agents Trade in Judgment Structures, Not Models

High-Value If, Low-Value Foreach: Why Agents Trade in Judgment Structures, Not Models

2
Comments
23 min read
Tackle High Token Usage with GraphRAG

Tackle High Token Usage with GraphRAG

1
Comments
4 min read
Why you still do not trust your AI's memory

Why you still do not trust your AI's memory

1
Comments
3 min read
How to build a production RAG pipeline in Python (without a vector database)

How to build a production RAG pipeline in Python (without a vector database)

1
Comments
5 min read
Como treinei uma IA de suporte com histórico real de atendimento: da conversa bruta ao RAG em produção

Como treinei uma IA de suporte com histórico real de atendimento: da conversa bruta ao RAG em produção

1
Comments 1
11 min read
Stop Burning Cash on Long-Context RAG: Ephemeral Prompt Caching with Spring AI and JTokkit

Stop Burning Cash on Long-Context RAG: Ephemeral Prompt Caching with Spring AI and JTokkit

Comments 1
2 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.