PHP is not where anyone expects GPU computing to happen. I wanted to find out how far it can go without a Python sidecar, so I built php-gpu-tensors, a native PHP extension for NVIDIA GPUs. This post walks through the interesting part: training a small neural network entirely in PHP, and what an optional kernel-fusion mode did to its speed.
What the extension gives you
CudaArray is a GPU tensor. Operations return GPU tensors, and data only crosses back to PHP when you call toArray():
use Cuda\CudaArray;
$input = new CudaArray([[1, 2], [3, 4]], 'float32');
$weights = CudaArray::ones([2, 2]);
$output = $input->add($weights)->multiply(2);
print_r($output->toArray()); // [[4, 6], [8, 10]]
You also get broadcasting, matmul(), reductions, Python-style slicing, .npy import, and the ability to compile your own CUDA C++ kernels at runtime with NVRTC.
The eager-mode problem
By default each operation is its own GPU call and produces its own result tensor. For a large model that's fine. For a small training step made of dozens of tiny operations, the per-operation overhead becomes the bottleneck, not the GPU's arithmetic.
Fusion: capture once, replay many times
Fusion is opt-in. You hand it a closure and example inputs; it runs the closure once with metadata-only placeholders, records the operations, and plans them:
use Cuda\Fusion;
$plan = Fusion::compile(
fn($a, $b, $c) => $a + $b * $c,
inputs: [$a, $b, $c]
);
$result = $plan->run($a, $b, $c); // no PHP callback code runs here
Elementwise operations are merged into generated CUDA kernels compiled with NVRTC. Matrix multiplications and reductions stay as "native boundaries" that use existing kernels, and everything is replayed on a private stream. For $a + $b * $c the plan is one fused kernel with no intermediate buffers, and getPlan() / getStats() show you exactly what the planner produced.
The training loop
fused.php trains a small ReLU classifier. The forward pass, a numerically stable softmax cross-entropy, the hand-written backward pass and clipped SGD are one closure, compiled once per batch shape and replayed with the updated parameters. Batches are uploaded once and stay on the GPU. Loss is only read back on reporting epochs.
There is no autograd; the backward pass is written explicitly, and a test checks one full step against eager execution, a CPU loss and finite-difference gradients.
Results
Hardware: an entry-level NVIDIA GeForce MX570 A (4 GB), PHP 8.3, driver 12.6, CUDA runtime 12.3.
I ran the same small MLP (hidden size 256, batch 512, 20 epochs) on a subset of PatchCamelyon (32,768 train and 4,096 test patches, 480 handcrafted colour features per patch), once in eager mode and once as a compiled Fusion plan:
| Mode | Time per step | Patches per second | Training time |
|---|---|---|---|
| Eager | 4.08 ms | 125,444 | 5.22 s |
| Fusion replay | 0.66 ms | 776,251 | 0.84 s |
That's about 6.2× faster, and the test metrics are identical in both modes (77.29% accuracy, AUC 0.853, same confusion matrix), which is the evidence that fusion doesn't change the math. The planner turned the 68 captured nodes of the training step into 11 fused kernels plus 9 native boundaries.
On UCI Optdigits (a 64-64-10 classifier, 1,000 epochs) the example reaches 97.25% test accuracy: 12,000 steps in 3.67 s, about 830k samples per second.
A caveat I want to be upfront about: these models are small, so most of the gain is removed per-operation overhead rather than raw GPU throughput. PatchCamelyon here is a performance demonstration, not a clinical model.
What it can't do (yet)
- It is a low-level library, not an ML framework: no autograd.
- NVIDIA GPUs and Linux only; no CPU fallback.
- Plans with matmul or reductions support async replay but not CUDA Graph yet.
- It's beta (
0.1.0-beta.4). I've validated it on an MX570 A with PHP 8.3 and earlier on an RTX A2000, so other GPUs are unexplored.
Try it
pie install lcmialichi/php-gpu-tensors:0.1.0-beta.4
php fused.php --epochs=200 --batch-size=256 --learning-rate=0.05 --no-save
It needs the CUDA Toolkit (including NVRTC) and an NVIDIA driver. The code, examples and benchmarks are on GitHub:
- https://github.com/lcmialichi/php-gpu-tensors
- https://github.com/lcmialichi/php-gpu-tensors-benchmarks
If you have an NVIDIA GPU, I'd especially like to hear whether ./run-tests.sh --require-gpu passes for you.
Top comments (0)