It's been more than a year since I last coded in C/C++, so I decided to start small by solving a few mathematical problems in C++. I've solved a couple so far.
Back in college, I was introduced to LLMs, and about a year later I came across Ollama. I was curious how long the GPU in my PC would take to process a request locally. So I went deeper and ran into GPUs, TPUs and NPUs. Then I noticed something unexpected: all of them are built to do one thing at massive scale, which is matrix multiplication.
That got me looking for ways to multiply matrices faster than the standard O(N³) approach. To my surprise, I found a 2022 paper from Google DeepMind called AlphaTensor, published before ChatGPT made AI a mainstream thing. The team used reinforcement learning to discover matrix multiplication algorithms that need fewer multiplications than the best ones humans had found. For example, it multiplied 4×4 matrices (in modular arithmetic) with 47 multiplications instead of Strassen's 49. To be honest, I was amazed at first. Then it hit me: they used an ML algorithm to speed up the very thing ML algorithms spend all their time doing. You've got to be kidding me, right?
Today, I decided to dig into how Google TPUs, NPUs and the mighty Nvidia GPUs are built on the inside. That's where I learned about GEMM, BLAS, systolic arrays, XLA, MXUs, neuromorphic chips, cuDNN and CUTLASS.
Starting today, I'm going to document everything I learn here, in public.
RIGHT TIME. RIGHT NOW.
Top comments (0)