DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
From Timeouts to Savings: How we optimized 24-page PDF parsing with Gemini & OpenRouter

From Timeouts to Savings: How we optimized 24-page PDF parsing with Gemini & OpenRouter

Comments
2 min read
Q4 KV Cache Fit 32K Context into 8GB VRAM — Only Math Broke

Q4 KV Cache Fit 32K Context into 8GB VRAM — Only Math Broke

Comments
8 min read
Nvidia Chips, AI Limitations, and Cybersecurity Shifts

Nvidia Chips, AI Limitations, and Cybersecurity Shifts

Comments
2 min read
Building a Mini Palantir: A Local Graph-RAG Engine with Ontology, Security, and Self-Evolution (Alpha)

Building a Mini Palantir: A Local Graph-RAG Engine with Ontology, Security, and Self-Evolution (Alpha)

1
Comments 1
6 min read
Why does paying more make your LLM reply faster?

Why does paying more make your LLM reply faster?

2
Comments 2
3 min read
We Tested 10 Untested LLMs on Agent Coding — The Results Are In

We Tested 10 Untested LLMs on Agent Coding — The Results Are In

3
Comments
3 min read
HBM4 Didn't Break the Memory Wall — It Just Moved It

HBM4 Didn't Break the Memory Wall — It Just Moved It

Comments
6 min read
Anthropic Just Released a Model So Dangerous They Gave It to Only Security Researchers

Anthropic Just Released a Model So Dangerous They Gave It to Only Security Researchers

Comments
2 min read
LLMKube Now Deploys Any Inference Engine, Not Just llama.cpp

LLMKube Now Deploys Any Inference Engine, Not Just llama.cpp

Comments
3 min read
80% of RAG Failures Start Here (And It's Not the LLM)

80% of RAG Failures Start Here (And It's Not the LLM)

4
Comments
2 min read
Anthropic caught its AI agent blackmailing to survive — here's how it's fixing it

Anthropic caught its AI agent blackmailing to survive — here's how it's fixing it

Comments
3 min read
Your text file is the prompt now: LLM's shebang trick

Your text file is the prompt now: LLM's shebang trick

2
Comments
3 min read
AI Agents for Enterprise Data Analytics: From Chat Interfaces to Reliable Execution

AI Agents for Enterprise Data Analytics: From Chat Interfaces to Reliable Execution

5
Comments
3 min read
Running Just One LLM on 8GB VRAM Is a Waste

Running Just One LLM on 8GB VRAM Is a Waste

Comments
8 min read
Solving the LLM Black Box Problem with Structured Reasoning

Solving the LLM Black Box Problem with Structured Reasoning

4
Comments 4
4 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.