DEV Community

#cuda

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Running the 510 GB DeepSeek-V4.1-Flash on an 8 GB GPU — and three bugs that never raise an error

Running the 510 GB DeepSeek-V4.1-Flash on an 8 GB GPU — and three bugs that never raise an error

Comments
8 min read
Running a 133 GB MoE model on an 8 GB GPU at 11 tokens/s by streaming experts from NVMe

Running a 133 GB MoE model on an 8 GB GPU at 11 tokens/s by streaming experts from NVMe

Comments
6 min read
Gemma 4 at Over 70 Tokens/s on a 2021 Laptop's 4 GB GPU: The Live Demo, Step by Step

Gemma 4 at Over 70 Tokens/s on a 2021 Laptop's 4 GB GPU: The Live Demo, Step by Step

9
Comments
12 min read
Gemma 4 on a Tesla T4, Part 3: Int4 Embeddings Serve E2B in 2.86 GiB at 2.30x bf16

Gemma 4 on a Tesla T4, Part 3: Int4 Embeddings Serve E2B in 2.86 GiB at 2.30x bf16

9
Comments
11 min read
How Fast Can a 421M-Parameter Decision Model Run? I Benchmarked Laya Across NVIDIA GPUs

How Fast Can a 421M-Parameter Decision Model Run? I Benchmarked Laya Across NVIDIA GPUs

Comments 1
6 min read
Exploring Result Visibility of Fixed-Latency Instructions on the SM120 Architecture

Exploring Result Visibility of Fixed-Latency Instructions on the SM120 Architecture

1
Comments
10 min read
Nvidia Just Let Rust Into CUDA. Here's Why That's a Bigger Deal Than It Sounds

Nvidia Just Let Rust Into CUDA. Here's Why That's a Bigger Deal Than It Sounds

1
Comments
5 min read
CUDA Rust: two native tracks, not a wrapper over C++

CUDA Rust: two native tracks, not a wrapper over C++

1
Comments
5 min read
A 4 GB Laptop GPU vs a 6-Core CPU on Gemma 4, Re-Measured in ABBA Order: 4.1x

A 4 GB Laptop GPU vs a 6-Core CPU on Gemma 4, Re-Measured in ABBA Order: 4.1x

10
Comments 2
11 min read
Reverse-Engineering NVIDIA: Modifying a CUDA binary

Reverse-Engineering NVIDIA: Modifying a CUDA binary

Comments
4 min read
Gemma 4 on a Tesla T4, Part 2: The Minimum GCE VM and a Script to Drive It

Gemma 4 on a Tesla T4, Part 2: The Minimum GCE VM and a Script to Drive It

13
Comments
13 min read
Finding a Random Island with Geometry and CUDA

Finding a Random Island with Geometry and CUDA

Comments
4 min read
Gemma 4 on a Tesla T4: QAT Weights Decode 1.79x Faster Than bf16

Gemma 4 on a Tesla T4: QAT Weights Decode 1.79x Faster Than bf16

8
Comments 1
9 min read
Same nvJPEG2000, different numbers: timer boundaries and frames in flight

Same nvJPEG2000, different numbers: timer boundaries and frames in flight

Comments
13 min read
Qwen3.8 27B at 256K: 50 TPS on a 24 GB GPU

Qwen3.8 27B at 256K: 50 TPS on a 24 GB GPU

1
Comments 1
11 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.