DEV Community

#llm

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
Stream LLM responses in a voice pipeline: Tool calling, structured outputs, and real-time actions

Stream LLM responses in a voice pipeline: Tool calling, structured outputs, and real-time actions

Comments
10 min read
Your AI agent remembers what sounds related, not what worked

Your AI agent remembers what sounds related, not what worked

2
Comments 13
5 min read
GPT-5.5 Pro เทียบกับ Instant: คุ้มค่าไหมเมื่อราคาต่าง 6 เท่า

GPT-5.5 Pro เทียบกับ Instant: คุ้มค่าไหมเมื่อราคาต่าง 6 เท่า

Comments
6 min read
AI Agents vs. Traditional Automation: When to Use Each

AI Agents vs. Traditional Automation: When to Use Each

Comments
13 min read
The Softmax Bottleneck: Why Making LLMs Bigger Doesn't Always Make Them Smarter

The Softmax Bottleneck: Why Making LLMs Bigger Doesn't Always Make Them Smarter

1
Comments
4 min read
We Built a 'Grovel Index' to Measure LLM Sycophancy —Here's What We Found

We Built a 'Grovel Index' to Measure LLM Sycophancy —Here's What We Found

2
Comments 7
5 min read
Counterintuitive: WSL2 + vllm cannot fit Qwen2.5-7B-1M on 6GB VRAM where Windows transformers can

Counterintuitive: WSL2 + vllm cannot fit Qwen2.5-7B-1M on 6GB VRAM where Windows transformers can

Comments
2 min read
Run GLM-5.2 Locally: The Open Model Nobody Can Ban

Run GLM-5.2 Locally: The Open Model Nobody Can Ban

20
Comments
10 min read
Why I Chose Free AI Models Over GPT-4 for Code Generation (And What Happened)

Why I Chose Free AI Models Over GPT-4 for Code Generation (And What Happened)

1
Comments
5 min read
Building a domain-specific LLM evaluation set from scratch

Building a domain-specific LLM evaluation set from scratch

1
Comments
8 min read
The Rule Held. The Boundary Moved Up. AI Memory Judgment, CLAIM-31: Verified Carryover Across Closes

The Rule Held. The Boundary Moved Up. AI Memory Judgment, CLAIM-31: Verified Carryover Across Closes

3
Comments
9 min read
Model Showdown Round 4: Opus vs Qwen — Writers, Not Coders

Model Showdown Round 4: Opus vs Qwen — Writers, Not Coders

Comments
10 min read
SubQ Model: Can Subquadratic Make Long-Context AI More Efficient?

SubQ Model: Can Subquadratic Make Long-Context AI More Efficient?

1
Comments
9 min read
The most underrated feature in a website AI assistant: saying "I don't know"

The most underrated feature in a website AI assistant: saying "I don't know"

Comments
2 min read
We burned 136 million tokens running an autonomous agent studio. Here's how we cut the bill ~90%.

We burned 136 million tokens running an autonomous agent studio. Here's how we cut the bill ~90%.

Comments
5 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.