DEV Community

Karthik Unnikrishnan
Karthik Unnikrishnan

Posted on

From Naive CUDA to Performance Engineering: My First GPU Matmul Journey

I have been diving deep into GPU architecture recently—balancing hands-on coding with studying Programming Massively Parallel Processors (PMPP)—and I am thrilled to share that I just earned my AMD Certified Associate credential along the way. To truly cement this architectural knowledge, I decided to stop just reading about kernels and actually build a complete GPU performance-engineering project from scratch. I chose matrix multiplication as the starting point because it exposes almost every critical aspect of GPU programming, from memory hierarchy to thread cooperation.

The journey began with a naive CUDA implementation. The basic idea was simple: assign one CUDA thread the responsibility of calculating exactly one element of the output matrix. Each thread determines its row and column based on its block and thread indices, and then performs the entire dot product independently. This gave me a functional, mathematically correct baseline. Having a baseline is extremely important because optimization is meaningless if you have nothing to compare your improvements against.

However, the naive implementation immediately highlights a fundamental GPU bottleneck: the memory wall. When several neighboring threads calculate their respective elements, they all need to fetch the same rows and columns from global memory repeatedly. Global memory has high latency and finite bandwidth, meaning these redundant data fetches cripple the kernel's actual compute potential. The solution, which aligns directly with the architectural concepts I have been studying, is to exploit data reuse.

To solve the memory bottleneck, I implemented a tiled matrix multiplication kernel using shared memory. Instead of every thread independently pulling data from global memory, threads within a block cooperate to load small portions—or "tiles"—of the input matrices into the SM's ultra-fast shared memory. They synchronize to ensure the tile is fully loaded, and then reuse those values for multiple calculations. I also made sure this wasn't just a toy implementation; I added proper boundary checks and zero-padding in shared memory so the kernel accurately processes arbitrary matrix dimensions (like 127 × 131) rather than just clean multiples of 16.

Writing the kernel was only half the project; measuring it accurately was the other. I refactored the codebase to cleanly separate the core CUDA kernels from the testing infrastructure. Relying on CPU wall-clock time is a terrible way to measure GPU performance, so I built a dedicated benchmark harness using CUDA events. The benchmark warms up the GPU, runs 100 iterations of the kernel, and records the GPU-side elapsed time to calculate average latency and GFLOP/s (floating-point operations per second). I integrated all of this into a strict GitHub engineering workflow—handling feature branches, pull requests, and correctness tests just like a production codebase.

The rigorous benchmarking paid off immediately. Testing a 1024 × 1024 × 1024 matrix on a Colab GPU, the naive kernel hit a latency of 4.660 ms (about 460 GFLOP/s). The tiled shared-memory kernel dropped that latency to 2.429 ms (884 GFLOP/s). That is a measured 1.92× speedup purely from optimizing how data moves through the GPU's memory hierarchy.

The most important shift in this project is that I am no longer asking, "Can I write a CUDA kernel?" I am now asking, "Why does this kernel perform this way, and what evidence tells me what to change?" I am intentionally avoiding the temptation to blindly throw warp shuffles or tensor cores at the code. My next steps are driven by data: sweeping matrix sizes from 512 to 4096 to understand scaling behavior, running tile-size experiments, and using profiling tools for roofline analysis. Once the foundational performance methodology is locked in, this pipeline will pave the way for building custom Softmax, LayerNorm, and eventually, a fused FlashAttention implementation.

GitHub logo Karthik-Unni / cuda-attention-lab

From-scratch CUDA kernels and GPU optimization experiments, progressing toward optimized Transformer attention.

CUDA Attention Lab

A from-scratch CUDA kernel optimization laboratory focused on understanding GPU architecture, memory hierarchy, parallel algorithms, and the techniques used to build efficient Transformer attention kernels.

The project starts with simple CUDA kernels and progressively transforms them into optimized implementations through measurement, profiling, and hardware-aware optimization.


Why this project?

Modern GPU performance is not achieved simply by writing a correct parallel algorithm.

A kernel can be mathematically correct and still perform poorly because of:

  • inefficient global-memory access
  • insufficient data reuse
  • poor memory coalescing
  • excessive synchronization
  • register pressure
  • shared-memory usage
  • low occupancy
  • instruction dependencies
  • memory bandwidth limitations
  • compute throughput limitations

This project is an experimental study of these problems.

No optimization is accepted simply because it "looks faster".


Project progression

1. CUDA Fundamentals

Learn the CUDA execution model:

  • threads
  • warps
  • blocks
  • grids
  • SMs
  • global memory
  • shared memory
  • registers
  • kernel launches
  • host/device memory transfers

2. Matrix Multiplication

Matrix multiplication is…

Top comments (0)