Dust: The Radical Idea of Training AI Without Backpropagation
What if everything we know about training neural networks is wrong?
For decades, deep learning has been built on one algorithm: backpropagation. It's the engine behind every transformer, every LLM, every image generator. But what if there's a simpler, more scalable alternative that we've been ignoring because backprop was "good enough"?
Enter Dust — a new zeroth-order optimization method that just achieved something remarkable: competitive pretraining performance without ever computing a gradient.
The Backprop Bottleneck
Backpropagation is elegant. It computes exactly how much each parameter contributed to the error, then adjusts accordingly. But it has fundamental constraints:
- Differentiability required — Your architecture must be differentiable end-to-end. This eliminates entire classes of models.
- Memory hungry — You must store all intermediate activations for the backward pass. This limits batch sizes and context lengths.
- Hardware constraints — GPUs are optimized for matrix multiplication and gradient computation. This locks us into specific hardware designs.
Backprop is an inductive bias — a helpful assumption that makes learning efficient in low-compute regimes. But as Sutton's "bitter lesson" teaches us: general methods that scale with compute eventually win.
How Dust Works
Dust replaces gradient computation with activation-space perturbation. Here's the core insight:
Instead of computing how each parameter affects the loss (backprop), Dust randomly perturbs the activations at each layer and observes which perturbations improve performance. Each token in a forward pass acts as an independent "population member" being evaluated in parallel.
Think of it like this:
- Backprop: Carefully calculate the exact slope of the loss landscape, then step downhill
- Dust: Take a thousand tiny random steps, keep the ones that go downhill, average them into a direction
The surprising finding: at sufficient scale, Dust's gradient estimates align closely with backprop's. And sometimes, they even exceed backprop performance.
The Results That Matter
1. Competitive with Backprop
At large population sizes, Dust matches backprop performance on pretraining tasks. This alone is shocking — zeroth-order methods were widely believed not to scale.
2. Larger Models Are MORE Efficient
Counterintuitively, bigger models are more population-efficient, not less. A 243M-parameter model outperforms a 120× smaller model at most population sizes. This is the opposite of what most researchers expected.
3. Massive Efficiency Gains
Dust is 10³ to 10⁴ times more efficient than weight-space evolutionary strategies. This makes it practical, not just theoretical.
4. Alignment Persists
Dust's gradient estimates stay well-aligned with backprop at every scale tested, up to 1B tokens. This suggests the method will continue working as models grow.
Why This Could Change Everything
Non-Differentiable Hardware
Dust could enable training on:
- Neuromorphic chips (like Intel's Loihi)
- Analog computing (which is inherently non-differentiable)
- Optical neural networks (light-based computation)
- Biological substrates (if we ever get there)
Memory Efficiency
No backward pass means no need to store activations. This could:
- Enable much larger batch sizes
- Allow training on memory-constrained devices
- Reduce the memory wall that's limiting context lengths
Biological Plausibility
Brains don't do backprop. They use local learning rules and Hebbian plasticity. Dust is closer to how biological learning might actually work — random perturbation + selection.
Architecture Freedom
Differentiability constrains architecture design. Dust removes this constraint, potentially unlocking entirely new model designs that are currently impossible to train.
The Philosophical Shift
There's something profound here. Backprop is a clever shortcut — it uses calculus to efficiently compute exact gradients. But shortcuts often become limitations.
Dust represents a return to first principles: if you have enough compute, maybe you don't need clever shortcuts. Maybe brute-force search + selection is enough.
This is the same lesson as AlphaGo Zero: bootstrapping on human knowledge helps initially, but pure self-play with enough compute eventually wins.
What This Means for AI Development
If Dust scales (and the early evidence suggests it will), we could see:
- Training on edge devices — Your phone could train models locally
- New hardware paradigms — Neuromorphic and analog AI accelerators become viable
- More efficient research — Faster experimentation with non-standard architectures
- Biological AI — A step toward AI systems that learn more like brains
We're not there yet — Dust is still early research. But the fact that it works at all challenges a fundamental assumption that's guided AI research for 40 years.
The Bottom Line
Dust is a reminder that even our most fundamental assumptions should be questioned. Backprop has been the foundation of deep learning since the 1980s. Now, in 2026, someone finally asked: "Is this actually the best way to do this?"
The answer might be no. And that opens doors we didn't even know existed.
What do you think? Could zeroth-order optimization replace backprop? Or will they coexist for different use cases? Let me know in the comments.
References:
- Dust paper: https://qlabs.sh/research/dust
- Sutton's Bitter Lesson: http://www.incompleteideas.net/IncIdeas/BitterLesson.html
- Code: https://github.com/qlabs-eng/dust
Top comments (0)