DEV Community

#gpu

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
A $22,000 GPU bill with a very quiet night shift

A $22,000 GPU bill with a very quiet night shift

2
Comments
1 min read
"We're not leaving: hunting B300s on demand across four suppliers (live x402 results)"

"We're not leaving: hunting B300s on demand across four suppliers (live x402 results)"

5
Comments
3 min read
"Four suppliers, real USDC, no accounts: the full results of our x402 GPU test"

"Four suppliers, real USDC, no accounts: the full results of our x402 GPU test"

5
Comments
3 min read
"We let an AI agent pay for GPUs with USDC. Here is what happened (2 of 3 delivered, 1 auto-refunded)"

"We let an AI agent pay for GPUs with USDC. Here is what happened (2 of 3 delivered, 1 auto-refunded)"

5
Comments
2 min read
I Compared 5 GPU Clouds for LLM Inference in 2026 — Here's What I Found

I Compared 5 GPU Clouds for LLM Inference in 2026 — Here's What I Found

Comments
2 min read
The FP8 trap: my GPU bill dropped 47% because the model was printing "!!!!!!"

The FP8 trap: my GPU bill dropped 47% because the model was printing "!!!!!!"

Comments
3 min read
⚡️One RTX 4090, 100 Trillion Tokens/Second – The Future of AI is in Your Desktop!

⚡️One RTX 4090, 100 Trillion Tokens/Second – The Future of AI is in Your Desktop!

Comments
6 min read
Running 100B+ MoE Models on a Single RTX 4090: A Practical Guide to Expert Offloading with llama.cpp

Running 100B+ MoE Models on a Single RTX 4090: A Practical Guide to Expert Offloading with llama.cpp

Comments
8 min read
Backends 101: Choosing the Right Measurement Surface

Backends 101: Choosing the Right Measurement Surface

Comments
3 min read
Running 100B+ MoE LLMs on a Single RTX 4090: A Practical Guide to Expert Offloading with llama.cpp

Running 100B+ MoE LLMs on a Single RTX 4090: A Practical Guide to Expert Offloading with llama.cpp

Comments
7 min read
Kubernetes SR-IOV Multi-Rail GPU Networking: Missing Route Configuration for Cross-Rail Communication

Kubernetes SR-IOV Multi-Rail GPU Networking: Missing Route Configuration for Cross-Rail Communication

Comments
11 min read
Inside XCP: How We Turned Xbox Memory and GPU Compute Into a Verifiable Worker

Inside XCP: How We Turned Xbox Memory and GPU Compute Into a Verifiable Worker

18
Comments
8 min read
Vast.ai CLI 101: Finding a GPU Offer You Can Actually Use

Vast.ai CLI 101: Finding a GPU Offer You Can Actually Use

Comments
7 min read
From Naive CUDA to Performance Engineering: My First GPU Matmul Journey

From Naive CUDA to Performance Engineering: My First GPU Matmul Journey

Comments
3 min read
Your Apple Silicon GPU Loses to One CPU Core Until a Million Rows. I Measured 111 Operations, Then Rebuilt ArrowMetal 0.2.0 Around the Answer

Your Apple Silicon GPU Loses to One CPU Core Until a Million Rows. I Measured 111 Operations, Then Rebuilt ArrowMetal 0.2.0 Around the Answer

Comments
10 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.