DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Counterintuitive: WSL2 + vllm cannot fit Qwen2.5-7B-1M on 6GB VRAM where Windows transformers can

Counterintuitive: WSL2 + vllm cannot fit Qwen2.5-7B-1M on 6GB VRAM where Windows transformers can

Comments
2 min read
Run GLM-5.2 Locally: The Open Model Nobody Can Ban

Run GLM-5.2 Locally: The Open Model Nobody Can Ban

20
Comments
10 min read
Why I Chose Free AI Models Over GPT-4 for Code Generation (And What Happened)

Why I Chose Free AI Models Over GPT-4 for Code Generation (And What Happened)

1
Comments
5 min read
Building a domain-specific LLM evaluation set from scratch

Building a domain-specific LLM evaluation set from scratch

1
Comments
8 min read
Model Showdown Round 4: Opus vs Qwen — Writers, Not Coders

Model Showdown Round 4: Opus vs Qwen — Writers, Not Coders

Comments
10 min read
SubQ Model: Can Subquadratic Make Long-Context AI More Efficient?

SubQ Model: Can Subquadratic Make Long-Context AI More Efficient?

1
Comments
9 min read
The Rule Held. The Boundary Moved Up. AI Memory Judgment, CLAIM-31: Verified Carryover Across Closes

The Rule Held. The Boundary Moved Up. AI Memory Judgment, CLAIM-31: Verified Carryover Across Closes

3
Comments
9 min read
The most underrated feature in a website AI assistant: saying "I don't know"

The most underrated feature in a website AI assistant: saying "I don't know"

Comments
2 min read
We burned 136 million tokens running an autonomous agent studio. Here's how we cut the bill ~90%.

We burned 136 million tokens running an autonomous agent studio. Here's how we cut the bill ~90%.

Comments
5 min read
LLM-as-Judge Is Three Decisions

LLM-as-Judge Is Three Decisions

Comments 1
6 min read
We Built a Self-Hosted AI Platform That Runs 100% on Your Hardware — Introducing local-ai.run

We Built a Self-Hosted AI Platform That Runs 100% on Your Hardware — Introducing local-ai.run

1
Comments
5 min read
6 lessons on testing AI features

6 lessons on testing AI features

1
Comments 8
8 min read
TokenSpeed and the Quiet Race to Make LLM Inference Boring

TokenSpeed and the Quiet Race to Make LLM Inference Boring

1
Comments 1
5 min read
🔬 Direction 1 closure on JAMES — when the hypothesis fails but the data turns "7-tier monotonic natural-stop gradient"

Gemma 4 Challenge: Write about Gemma 4 Submission

🔬 Direction 1 closure on JAMES — when the hypothesis fails but the data turns "7-tier monotonic natural-stop gradient"

1
Comments
2 min read
ExLlamaV3 Updates, Unsloth Qwen GGUFs & Phi3 Autonomous Bridge

ExLlamaV3 Updates, Unsloth Qwen GGUFs & Phi3 Autonomous Bridge

Comments 1
3 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.