Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
llminference
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
VKAE: VIDRAFT's Inference Engine Hits 23 GPU Speedup and ~10K Tokens/sec on a Single Nvidia B200
AI OpenFree
AI OpenFree
AI OpenFree
Follow
Aug 7
VKAE: VIDRAFT's Inference Engine Hits 23 GPU Speedup and ~10K Tokens/sec on a Single Nvidia B200
#
llminference
#
gpuacceleration
#
vkae
#
nvidiab200
5
 reactions
Comments
Add Comment
4 min read
AI Inference at the Edge: Running Real-Time LLMs in Kubernetes Without a GPU Farm
The Cyber Sidekick
The Cyber Sidekick
The Cyber Sidekick
Follow
Jun 18
AI Inference at the Edge: Running Real-Time LLMs in Kubernetes Without a GPU Farm
#
edgeai
#
kubernetes
#
llminference
#
vllm
Comments
Add Comment
3 min read
Qwen 3.6 enable_thinking — The MoE Pitfall That Broke My Agent JSON Parsing
SleepyQuant
SleepyQuant
SleepyQuant
Follow
May 18
Qwen 3.6 enable_thinking — The MoE Pitfall That Broke My Agent JSON Parsing
#
qwen
#
mlx
#
localai
#
llminference
Comments
Add Comment
5 min read
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account