Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
llm
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
Ollama 0.30 GPU Boost: Faster local Qwen inference on NVIDIA
EveryLocalAI
EveryLocalAI
EveryLocalAI
Follow
Jun 10
Ollama 0.30 GPU Boost: Faster local Qwen inference on NVIDIA
#
ai
#
llm
#
performance
#
tutorial
Comments
Add Comment
1 min read
Making a fleet of self-hosted LLM agents trustworthy
Christopher Maher
Christopher Maher
Christopher Maher
Follow
Jun 14
Making a fleet of self-hosted LLM agents trustworthy
#
ai
#
llm
#
kubernetes
#
opensource
1
 reaction
Comments
Add Comment
6 min read
Local LLM vs Claude: Benchmarking qwen3-coder:30b as a Production Agent Backend
Arsen Apostolov
Arsen Apostolov
Arsen Apostolov
Follow
Jul 3
Local LLM vs Claude: Benchmarking qwen3-coder:30b as a Production Agent Backend
#
llm
#
homelab
#
opensource
#
ai
Comments
1
 comment
5 min read
Why We Added Rate Limits Between AI Agents
Karan Padhiyar
Karan Padhiyar
Karan Padhiyar
Follow
Jun 10
Why We Added Rate Limits Between AI Agents
#
ai
#
infrastructure
#
llm
#
brainpackai
Comments
Add Comment
3 min read
Your schema validation passes and the agent still picks the wrong tool. The bug is semantic.
James O'Connor
James O'Connor
James O'Connor
Follow
Jun 10
Your schema validation passes and the agent still picks the wrong tool. The bug is semantic.
#
ai
#
llm
#
python
#
programming
Comments
Add Comment
2 min read
My Agent's Memory File Wasn't Wrong. It Was Just Six Weeks Stale.
Enjoy Kumawat
Enjoy Kumawat
Enjoy Kumawat
Follow
Jul 14
My Agent's Memory File Wasn't Wrong. It Was Just Six Weeks Stale.
#
ai
#
llm
#
agents
#
productivity
Comments
Add Comment
4 min read
LLM-as-Judge Shouldn't Aggregate Scores: Binary Checks as Evidence, One Holistic Verdict
shimo4228
shimo4228
shimo4228
Follow
Jul 14
LLM-as-Judge Shouldn't Aggregate Scores: Binary Checks as Evidence, One Holistic Verdict
#
llm
#
promptengineering
#
evaluation
#
claudecode
Comments
Add Comment
12 min read
Building an MCP Server in Python — Architecture, FastMCP, and Production Code
Piotrek Karasinski
Piotrek Karasinski
Piotrek Karasinski
Follow
Jul 3
Building an MCP Server in Python — Architecture, FastMCP, and Production Code
#
ai
#
python
#
architecture
#
llm
2
 reactions
Comments
Add Comment
8 min read
The Prefill Wall: Why MTP's 2 Barely Moves Long-Context Latency (Qwen3.6-27B, RTX 3090)
byeongsoo kang
byeongsoo kang
byeongsoo kang
Follow
Jun 10
The Prefill Wall: Why MTP's 2 Barely Moves Long-Context Latency (Qwen3.6-27B, RTX 3090)
#
llm
#
performance
#
machinelearning
#
rag
Comments
Add Comment
4 min read
My Home AI's First Reply Took Four Minutes. Now It Takes Eleven Seconds.
Nova
Nova
Nova
Follow
Jul 14
My Home AI's First Reply Took Four Minutes. Now It Takes Eleven Seconds.
#
ai
#
llm
#
selfhosted
#
devops
2
 reactions
Comments
Add Comment
4 min read
7 things I learned trying to stop LLM API bills from silently exploding
kimbeomgyu
kimbeomgyu
kimbeomgyu
Follow
Jul 12
7 things I learned trying to stop LLM API bills from silently exploding
#
ai
#
webdev
#
llm
#
node
4
 reactions
Comments
11
 comments
3 min read
The OWASP Agentic Top 10, explained for practitioners
Brenn Hill
Brenn Hill
Brenn Hill
Follow
Jul 14
The OWASP Agentic Top 10, explained for practitioners
#
security
#
ai
#
llm
#
owasp
1
 reaction
Comments
Add Comment
4 min read
How to Build AI Agents from Scratch with FastAPI
Ayush Kumar
Ayush Kumar
Ayush Kumar
Follow
Jul 14
How to Build AI Agents from Scratch with FastAPI
#
aiagents
#
python
#
llm
#
tutorial
Comments
Add Comment
8 min read
I Built an AI Agent That Writes Tests, Finds Bugs, and Opens PRs — Autonomously
Shivam Dhakad
Shivam Dhakad
Shivam Dhakad
Follow
Jun 11
I Built an AI Agent That Writes Tests, Finds Bugs, and Opens PRs — Autonomously
#
java
#
springboot
#
springai
#
llm
Comments
1
 comment
5 min read
Kubernetes in LLMOps (Part 2): GPU Efficiency, Cost Engineering, and Real-World Failure Modes
Mohammad Heydari
Mohammad Heydari
Mohammad Heydari
Follow
Jun 23
Kubernetes in LLMOps (Part 2): GPU Efficiency, Cost Engineering, and Real-World Failure Modes
#
infrastructure
#
kubernetes
#
llm
#
performance
1
 reaction
Comments
1
 comment
5 min read
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account