<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Lucas Mialichi</title>
    <description>The latest articles on DEV Community by Lucas Mialichi (@lcmialichi).</description>
    <link>https://dev.to/lcmialichi</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4161458%2Fee228260-1cb7-428e-a335-18448341acaa.jpg</url>
      <title>DEV Community: Lucas Mialichi</title>
      <link>https://dev.to/lcmialichi</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/lcmialichi"/>
    <language>en</language>
    <item>
      <title>Training a neural network on the GPU in pure PHP (and what kernel fusion did for it)</title>
      <dc:creator>Lucas Mialichi</dc:creator>
      <pubDate>Sun, 04 Oct 2026 11:17:20 +0000</pubDate>
      <link>https://dev.to/lcmialichi/training-a-neural-network-on-the-gpu-in-pure-php-and-what-kernel-fusion-did-for-it-e12</link>
      <guid>https://dev.to/lcmialichi/training-a-neural-network-on-the-gpu-in-pure-php-and-what-kernel-fusion-did-for-it-e12</guid>
      <description>&lt;p&gt;PHP is not where anyone expects GPU computing to happen. I wanted to find out how far it can go without a Python sidecar, so I built &lt;strong&gt;php-gpu-tensors&lt;/strong&gt;, a native PHP extension for NVIDIA GPUs. This post walks through the interesting part: training a small neural network entirely in PHP, and what an optional kernel-fusion mode did to its speed.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the extension gives you
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;CudaArray&lt;/code&gt; is a GPU tensor. Operations return GPU tensors, and data only crosses back to PHP when you call &lt;code&gt;toArray()&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight php"&gt;&lt;code&gt;&lt;span class="kn"&gt;use&lt;/span&gt; &lt;span class="nc"&gt;Cuda\CudaArray&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="nv"&gt;$input&lt;/span&gt;   &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;CudaArray&lt;/span&gt;&lt;span class="p"&gt;([[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;]],&lt;/span&gt; &lt;span class="s1"&gt;'float32'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="nv"&gt;$weights&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;CudaArray&lt;/span&gt;&lt;span class="o"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;ones&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;]);&lt;/span&gt;
&lt;span class="nv"&gt;$output&lt;/span&gt;  &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nv"&gt;$input&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;$weights&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="nf"&gt;multiply&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="nb"&gt;print_r&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;$output&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="nf"&gt;toArray&lt;/span&gt;&lt;span class="p"&gt;());&lt;/span&gt; &lt;span class="c1"&gt;// [[4, 6], [8, 10]]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You also get broadcasting, &lt;code&gt;matmul()&lt;/code&gt;, reductions, Python-style slicing, &lt;code&gt;.npy&lt;/code&gt; import, and the ability to compile your own CUDA C++ kernels at runtime with NVRTC.&lt;/p&gt;

&lt;h2&gt;
  
  
  The eager-mode problem
&lt;/h2&gt;

&lt;p&gt;By default each operation is its own GPU call and produces its own result tensor. For a large model that's fine. For a small training step made of dozens of tiny operations, the per-operation overhead becomes the bottleneck, not the GPU's arithmetic.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fusion: capture once, replay many times
&lt;/h2&gt;

&lt;p&gt;Fusion is opt-in. You hand it a closure and example inputs; it runs the closure once with metadata-only placeholders, records the operations, and plans them:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight php"&gt;&lt;code&gt;&lt;span class="kn"&gt;use&lt;/span&gt; &lt;span class="nc"&gt;Cuda\Fusion&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="nv"&gt;$plan&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Fusion&lt;/span&gt;&lt;span class="o"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;compile&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="k"&gt;fn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;$a&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;$b&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;$c&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nv"&gt;$a&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nv"&gt;$b&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="nv"&gt;$c&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;inputs&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;$a&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;$b&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;$c&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="nv"&gt;$result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nv"&gt;$plan&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;$a&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;$b&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;$c&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;   &lt;span class="c1"&gt;// no PHP callback code runs here&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Elementwise operations are merged into generated CUDA kernels compiled with NVRTC. Matrix multiplications and reductions stay as "native boundaries" that use existing kernels, and everything is replayed on a private stream. For &lt;code&gt;$a + $b * $c&lt;/code&gt; the plan is one fused kernel with no intermediate buffers, and &lt;code&gt;getPlan()&lt;/code&gt; / &lt;code&gt;getStats()&lt;/code&gt; show you exactly what the planner produced.&lt;/p&gt;

&lt;h2&gt;
  
  
  The training loop
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/lcmialichi/php-gpu-tensors/blob/main/fused.php" rel="noopener noreferrer"&gt;&lt;code&gt;fused.php&lt;/code&gt;&lt;/a&gt; trains a small ReLU classifier. The forward pass, a numerically stable softmax cross-entropy, the hand-written backward pass and clipped SGD are one closure, compiled once per batch shape and replayed with the updated parameters. Batches are uploaded once and stay on the GPU. Loss is only read back on reporting epochs.&lt;/p&gt;

&lt;p&gt;There is no autograd; the backward pass is written explicitly, and a test checks one full step against eager execution, a CPU loss and finite-difference gradients.&lt;/p&gt;

&lt;h2&gt;
  
  
  Results
&lt;/h2&gt;

&lt;p&gt;Hardware: an entry-level NVIDIA GeForce MX570 A (4 GB), PHP 8.3, driver 12.6, CUDA runtime 12.3.&lt;/p&gt;

&lt;p&gt;I ran the same small MLP (hidden size 256, batch 512, 20 epochs) on a subset of &lt;a href="https://github.com/basveeling/pcam" rel="noopener noreferrer"&gt;PatchCamelyon&lt;/a&gt; (32,768 train and 4,096 test patches, 480 handcrafted colour features per patch), once in eager mode and once as a compiled Fusion plan:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Mode&lt;/th&gt;
&lt;th&gt;Time per step&lt;/th&gt;
&lt;th&gt;Patches per second&lt;/th&gt;
&lt;th&gt;Training time&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Eager&lt;/td&gt;
&lt;td&gt;4.08 ms&lt;/td&gt;
&lt;td&gt;125,444&lt;/td&gt;
&lt;td&gt;5.22 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fusion replay&lt;/td&gt;
&lt;td&gt;0.66 ms&lt;/td&gt;
&lt;td&gt;776,251&lt;/td&gt;
&lt;td&gt;0.84 s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That's about 6.2× faster, and the test metrics are identical in both modes (77.29% accuracy, AUC 0.853, same confusion matrix), which is the evidence that fusion doesn't change the math. The planner turned the 68 captured nodes of the training step into 11 fused kernels plus 9 native boundaries.&lt;/p&gt;

&lt;p&gt;On UCI Optdigits (a 64-64-10 classifier, 1,000 epochs) the example reaches 97.25% test accuracy: 12,000 steps in 3.67 s, about 830k samples per second.&lt;/p&gt;

&lt;p&gt;A caveat I want to be upfront about: these models are small, so most of the gain is removed per-operation overhead rather than raw GPU throughput. PatchCamelyon here is a performance demonstration, not a clinical model.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it can't do (yet)
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;It is a low-level library, not an ML framework: no autograd.&lt;/li&gt;
&lt;li&gt;NVIDIA GPUs and Linux only; no CPU fallback.&lt;/li&gt;
&lt;li&gt;Plans with matmul or reductions support async replay but not CUDA Graph yet.&lt;/li&gt;
&lt;li&gt;It's beta (&lt;code&gt;0.1.0-beta.4&lt;/code&gt;). I've validated it on an MX570 A with PHP 8.3 and earlier on an RTX A2000, so other GPUs are unexplored.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pie &lt;span class="nb"&gt;install &lt;/span&gt;lcmialichi/php-gpu-tensors:0.1.0-beta.4
php fused.php &lt;span class="nt"&gt;--epochs&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;200 &lt;span class="nt"&gt;--batch-size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;256 &lt;span class="nt"&gt;--learning-rate&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;0.05 &lt;span class="nt"&gt;--no-save&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It needs the CUDA Toolkit (including NVRTC) and an NVIDIA driver. The code, examples and benchmarks are on GitHub:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/lcmialichi/php-gpu-tensors" rel="noopener noreferrer"&gt;https://github.com/lcmialichi/php-gpu-tensors&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/lcmialichi/php-gpu-tensors-benchmarks" rel="noopener noreferrer"&gt;https://github.com/lcmialichi/php-gpu-tensors-benchmarks&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you have an NVIDIA GPU, I'd especially like to hear whether &lt;code&gt;./run-tests.sh --require-gpu&lt;/code&gt; passes for you.&lt;/p&gt;

</description>
      <category>machinelearning</category>
      <category>performance</category>
      <category>php</category>
      <category>programming</category>
    </item>
  </channel>
</rss>
