DEV Community

#cuda

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
The sm_120 shared-memory cliff: why FP8 KV cache crashes vLLM on workstation Blackwell

The sm_120 shared-memory cliff: why FP8 KV cache crashes vLLM on workstation Blackwell

Comments
3 min read
The Cheapest CUDA GPU on AWS Has an Arm CPU — and You Probably Want the Intel One

The Cheapest CUDA GPU on AWS Has an Arm CPU — and You Probably Want the Intel One

3
Comments
11 min read
Accelerating Physical AI Workloads with NVIDIA CUDA and TensorRT

Accelerating Physical AI Workloads with NVIDIA CUDA and TensorRT

Comments
2 min read
I Built a CUDA Engine That Streams 744B Parameter AI Models on Consumer Hardware

I Built a CUDA Engine That Streams 744B Parameter AI Models on Consumer Hardware

2
Comments
4 min read
Qwen3-8B on workstation Blackwell: vLLM vs SGLang vs llama.cpp, plus an FP8 pass

Qwen3-8B on workstation Blackwell: vLLM vs SGLang vs llama.cpp, plus an FP8 pass

2
Comments 1
3 min read
Running Gemma 4 on EC2 G5g: Graviton2 AMD with NVIDIA GPU

Running Gemma 4 on EC2 G5g: Graviton2 AMD with NVIDIA GPU

13
Comments
9 min read
Breaking CUDA's Chains: Open-Source Alternatives for Cross-GPU Performance Optimization

Breaking CUDA's Chains: Open-Source Alternatives for Cross-GPU Performance Optimization

Comments
2 min read
Serving Gemma4 with Rust on vLLM 🦀

Serving Gemma4 with Rust on vLLM 🦀

10
Comments
10 min read
Running Gemma 4 on EC2 G5g: Graviton2 AMD with NVIDIA GPU

Running Gemma 4 on EC2 G5g: Graviton2 AMD with NVIDIA GPU

Comments
9 min read
RTX 5090 survival guide: sm_120, CUDA 12 and 13 side by side, and the xformers trap

RTX 5090 survival guide: sm_120, CUDA 12 and 13 side by side, and the xformers trap

1
Comments
7 min read
GPU-Accelerating MSCRED with CUDA, im2col, GEMM, and a Custom PyTorch Extension

GPU-Accelerating MSCRED with CUDA, im2col, GEMM, and a Custom PyTorch Extension

4
Comments 2
11 min read
Resurrecting Kepler: Getting Modern LLMs Running on a GTX 770 (Kernel 7.x)

Resurrecting Kepler: Getting Modern LLMs Running on a GTX 770 (Kernel 7.x)

1
Comments
4 min read
I Built Flash Attention From Scratch — Here's What Nobody Tells You About It

I Built Flash Attention From Scratch — Here's What Nobody Tells You About It

Comments
2 min read
One RTX 5090 vs a 12-GPU Cluster — Benchmarking a Decade of GPUs on the Same Go Proof

One RTX 5090 vs a 12-GPU Cluster — Benchmarking a Decade of GPUs on the Same Go Proof

Comments
4 min read
Learn CUDA and GPU programming without owning a GPU

Learn CUDA and GPU programming without owning a GPU

Comments
2 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.