I hypothesized that spiking neural networks would outperform CNNs on temporal tasks. I was wrong.
Background
My previous experiment showed SNNs match CNNs on static image classification (97.4% vs 97.8% on MNIST) but are 36x slower on conventional hardware. The obvious follow-up: test tasks where temporal dynamics should give SNNs a natural advantage.
The Tasks
Task 1: Sequential MNIST. Feed image rows one at a time (28 timesteps of 28 pixels). The model must accumulate spatial information across time steps. This should favor architectures with memory (LSTM, SNN membrane potential).
Task 2: Temporal spike patterns. Classify synthetic spike trains by their firing pattern: regular bursts, accelerating frequency, decelerating frequency, fast-then-slow, and bimodal inter-spike intervals. Pure temporal structure, no spatial component. This is the ideal SNN task.
Results
Sequential MNIST
| Architecture | Params | Accuracy | Train Time | Inference |
|---|---|---|---|---|
| CNN-1D | 11K | 93.15% | 17s | 0.016ms |
| LSTM | 25K | 92.95% | 59s | 0.146ms |
| SNN | 2.5K | 42.85% | 45s | 0.146ms |
SNN completely failed. 42.85% is barely above random for 10 classes.
Temporal Spike Patterns
| Architecture | Params | Accuracy | Train Time | Inference |
|---|---|---|---|---|
| CNN-1D | 4K | 99.75% | 9s | 0.023ms |
| SNN | 453 | 60.00% | 123s | 0.397ms |
| LSTM | 17K | 51.75% | 163s | 0.889ms |
CNN destroyed both temporal architectures. 99.75% vs 60% and 51.75%.
What Went Wrong With My Hypothesis
I assumed temporal tasks need temporal architectures. They don't. A 1D CNN treats a time series as a spatial signal and applies learned filters. Detecting "regular bursts every 20 steps" is a pattern detection problem, and convolutions are excellent at pattern detection regardless of whether the patterns are in space or time.
The SNN had too few parameters. 453 params (Task 2) vs 4,005 for CNN. The LIF neuron's linear transform plus membrane dynamics isn't expressive enough to learn complex temporal features. But adding more parameters defeats the SNN's efficiency argument.
Membrane decay forgets too fast. With beta=0.9, the membrane potential decays to 12% of its value after 20 timesteps. For patterns that span 100+ timesteps, the early information is gone. Higher beta (0.99) would retain more but makes the surrogate gradient optimization harder.
The One Thing SNNs DID Do Better
SNN beat LSTM on temporal patterns: 60% vs 51.75%, with 38x fewer parameters. Both are bad compared to CNN, but the SNN's LIF dynamics captured more temporal structure than LSTM's gating mechanism at this scale. The LSTM likely needs more hidden units and training time to converge on this task.
Revised Understanding
After two experiments:
- Static images: SNN matches CNN on accuracy, 36x slower
- Sequential images: SNN fails catastrophically (43% vs 93%)
- Temporal patterns: SNN loses to CNN-1D but beats LSTM at tiny scale
CNNs win because convolutions are universal pattern detectors. Whether the pattern is spatial (edges in an image) or temporal (bursts in a spike train), learned convolutional filters find them efficiently. The kernel slides across space or time identically.
SNNs are not better at temporal tasks on conventional hardware. Their advantage is energy efficiency on neuromorphic hardware, not architectural superiority on temporal data.
Code
python3 snn_temporal.py
github.com/turingrtss/vulndetect
Publishing negative results matters. The hypothesis was reasonable, the data said otherwise, and now I know something I did not know before.
Top comments (0)