<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Dijo Torres</title>
    <description>The latest articles on DEV Community by Dijo Torres (@dijo_torres).</description>
    <link>https://dev.to/dijo_torres</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4164650%2Ff05cd421-6a5e-4c54-a08f-3c5c191413fc.png</url>
      <title>DEV Community: Dijo Torres</title>
      <link>https://dev.to/dijo_torres</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/dijo_torres"/>
    <language>en</language>
    <item>
      <title>A Proper Kick-Off to My Embedded Journey!</title>
      <dc:creator>Dijo Torres</dc:creator>
      <pubDate>Mon, 05 Oct 2026 18:29:11 +0000</pubDate>
      <link>https://dev.to/dijo_torres/a-proper-kick-off-to-my-embedded-journey-563h</link>
      <guid>https://dev.to/dijo_torres/a-proper-kick-off-to-my-embedded-journey-563h</guid>
      <description>&lt;p&gt;It's been more than a year since I last coded in C/C++, so I decided to start small by solving a few mathematical problems in C++. I've solved a couple so far.&lt;/p&gt;

&lt;p&gt;Back in college, I was introduced to LLMs, and about a year later I came across Ollama. I was curious how long the GPU in my PC would take to process a request locally. So I went deeper and ran into GPUs, TPUs and NPUs. Then I noticed something unexpected: all of them are built to do one thing at massive scale, which is matrix multiplication.&lt;/p&gt;

&lt;p&gt;That got me looking for ways to multiply matrices faster than the standard O(N³) approach. To my surprise, I found a 2022 paper from Google DeepMind called AlphaTensor, published before ChatGPT made AI a mainstream thing. The team used reinforcement learning to discover matrix multiplication algorithms that need fewer multiplications than the best ones humans had found. For example, it multiplied 4×4 matrices (in modular arithmetic) with 47 multiplications instead of Strassen's 49. To be honest, I was amazed at first. Then it hit me: they used an ML algorithm to speed up the very thing ML algorithms spend all their time doing. You've got to be kidding me, right?&lt;/p&gt;

&lt;p&gt;Today, I decided to dig into how Google TPUs, NPUs and the mighty Nvidia GPUs are built on the inside. That's where I learned about GEMM, BLAS, systolic arrays, XLA, MXUs, neuromorphic chips, cuDNN and CUTLASS.&lt;/p&gt;

&lt;p&gt;Starting today, I'm going to document everything I learn here, in public.&lt;/p&gt;

&lt;p&gt;RIGHT TIME. RIGHT NOW.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>mojo</category>
      <category>compiling</category>
    </item>
  </channel>
</rss>
