How I Rediscovered Kepler's Third Law on a Phone, in 204 Lines
A minimal symbolic regression engine that runs on Termux, with no GPU and no cloud.
The problem with machine learning today
Most ML models are black boxes. You feed them data, they give you predictions, but they never tell you why.
For some tasks, that's fine. For others — physics, biology, finance — it's a problem. You don't want a prediction. You want a law.
That's what symbolic regression does. It doesn't fit parameters. It finds the equation itself.
It's a well-established field. Tools exist:
- PySR (Julia backend, state of the art)
- gplearn (Python, sklearn-compatible)
- FastSymbolicGP (2026, mobile preset)
- Eureqa (historical, closed source)
But they all share a common trait: they need a real computer. A GPU helps. A server is often required.
I wanted to see if the same idea could run on the smallest machine I had: my phone.
The constraint
My setup:
- Android phone
- Termux (a Linux environment for Android)
- Python 3.13
- NumPy, SymPy
- No GPU. No cloud. No server.
The goal was simple: build the smallest symbolic regression engine that fits on a phone.
Not the fastest. Not the most capable. The smallest.
What it does
The engine takes numeric data and discovers the underlying formula.
Two examples that worked:
| Target | Discovered | MSE | Time |
|---|---|---|---|
| y = x*sin(x) | y = x*sin(x) | 5.36e-33 | 18 ms |
| T = a^(3/2) | y = a**(3/2) | 4.77e-31 | 11 ms |
The second one is the Third Law of Kepler: the square of the orbital period is proportional to the cube of the semi-major axis.
The engine rediscovered it, from raw numbers, in 11 milliseconds, on an Android phone.
How it works
Three steps, in about 200 lines.
1. Build a library of candidate forms
For each input variable x, the engine generates: x, x^2, x^3, 1/x, 1/x^2, sqrt(x), sin(x), cos(x), exp(-x), log(x).
Then it generates all pairwise products: (a)*(b).
Total: around 55 candidate forms.
2. Sparse selection with OMP
Orthogonal Matching Pursuit picks the smallest subset of columns that best reconstructs the target.
Implemented in pure NumPy, in about 20 lines. It tries k = 1 through k = 6 and keeps the best MSE.
3. Symbolic simplification with SymPy
The selected terms are combined into a SymPy expression and simplified.
That's how a * sqrt(a) becomes a**(3/2).
What fails
Two targets didn't work.
Newton's law of gravitation (F = m1*m2/r^2): MSE 1.6.
Reason: the library only generates pairwise products. (m1*m2)*(1/r^2) is a triple product, and it isn't there.
Radioactive decay (N = exp(-t/2)): MSE 6e-6.
Reason: the library has exp(-t), but not exp(-t/2). Fractional exponents aren't generated.
Success rate: 2 out of 4 on known targets.
That's the expected behavior of a minimal engine. Not a bug. A design limit.
What this is — and what it isn't
It is: a pedagogical demonstration, a proof of feasibility, 204 lines of readable Python, a tool that fits on a phone.
It is not: production-ready, faster than PySR, a new scientific discovery, competitive with state-of-the-art tools.
If you want to solve a hard problem, use PySR. If you want to understand how symbolic regression works, this is a 200-line starting point.
Why it matters
Because the barrier to entry should be low.
You don't need a GPU cluster to study symbolic regression. You don't need a server to discover Kepler's third law.
You need a phone, an idea, and 200 lines of code.
Try it
pip install numpy sympy rich
python flash.py kepler
The full source is on GitHub: github.com/alpha0az1omega-sketch/adn-symbolic-regression
MIT license. Runs offline.
Written by alpha0az1omega. Built on Termux, Android. No GPU. No cloud.
Top comments (0)