DEV Community

turingrtss
turingrtss

Posted on

9 Neural Network Architectures on the Same Task. Here is What Won.

I ran 9 neural network architectures on the same dataset with the same budget. Here is what won and what surprised me.

Setup

Dataset: MNIST (20K training subset, 10K test)
Budget: 15 epochs, Adam 1e-3, same for everyone
Hardware: ARM64, 12GB RAM, 2 CPU cores, no GPU

Nine architectures: MLP, CNN, RNN, LSTM, GRU, Transformer, State Space (S4-style), KAN (Kolmogorov-Arnold Network), and Mixture of Experts.

Results

Architecture Params Accuracy Train Time Inference
CNN (LeNet) 54K 97.18% 676s 0.24ms
Transformer 19K 96.28% 516s 0.22ms
MLP 55K 94.71% 18s 0.01ms
Mixture of Experts 78K 94.30% 23s 0.01ms
LSTM 15K 93.83% 200s 0.28ms
GRU 15K 93.43% 147s 0.19ms
KAN 127K 87.65% 32s 0.06ms
Vanilla RNN 7K 76.17% 151s 0.14ms
State Space (S4) 4K 73.96% 215s 0.20ms

What I Learned

CNN still wins on image tasks. Not surprising for MNIST, but the margin matters. 97.18% vs the next best (Transformer at 96.28%) is a meaningful gap, and CNN got there with the most natural inductive bias for spatial data.

Transformer punches above its weight. Only 19K parameters (smallest after RNN and S4) but second highest accuracy. The attention mechanism captures row-to-row dependencies in the image effectively. Slowest to train though, because attention is O(n squared) on the 28-step sequence.

MLP is criminally underrated. 94.71% accuracy with 0.01ms inference. That is 24x faster than CNN for a 2.5% accuracy tradeoff. For any application where speed matters more than squeezing the last 2%, MLP wins.

KAN disappoints. The Kolmogorov-Arnold Network used 127K parameters (2.3x more than CNN) and only hit 87.65%. The learnable activation functions (approximated with RBF basis) are expressive in theory but hard to optimize. With 15 epochs it did not converge. More epochs might help, but the parameter efficiency is poor.

S4 needs more than a simplified implementation. My state space model is a crude approximation of the real S4/Mamba architecture. 73.96% is below what the architecture should achieve. This is an implementation gap, not an architecture gap.

Vanilla RNN confirms textbook knowledge. 76.17% on a 28-step sequence shows the vanishing gradient problem in action. LSTM and GRU fix this (93%+), which is exactly what they were designed for.

Mixture of Experts is efficient but not top-tier. 94.30% with two experts and a gating network. On a task this simple, specialization does not help much over a single larger network.

Speed vs Accuracy Tradeoff

MLP and MoE are 24-28x faster than CNN/Transformer at inference. On edge devices (Raspberry Pi, Jetson Nano, microcontrollers) where every millisecond matters, the simple architectures have a real deployment advantage.

MLP trained in 18 seconds. CNN took 676 seconds (37x longer). For rapid prototyping, the fast architectures let you iterate 37x more experiments in the same time.

Code

All implementations in a single file, pure PyTorch, no external dependencies:

python3 arch_comparison.py
Enter fullscreen mode Exit fullscreen mode

github.com/turingrtss/vulndetect


Next: spiking neural networks and a direct comparison of neuromorphic vs transformer architectures on the same task.

Top comments (1)

Some comments may only be visible to logged-in visitors. Sign in to view all comments.