This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass
What I Built
Paper Trials is a lab notebook for paper airplanes. It doesn't fold anything for you and it doesn't simulate flight. It tells you which one thing to change about your plane, predicts how far it will fly, and then asks you to close the tab and go outside.
The loop is:
- Plan. Pick a change, like "wing span 20 cm → 21 cm". Only one thing changes, so you'll know what made the difference.
- Predict. The app predicts the median distance with a range and saves it. It can't be edited after that, so you can't quietly move the goalposts once you've seen the result.
- Fly. Build it, go outside, throw it three times the same way, and measure where it lands. A tape is nice; counting steps works.
- Log. Come back, type in the distances, and add a note if something happened ("climbed, stalled, pulled left").
- Learn. The model learns from every flight, and the notebook suggests the next change.
The screen is only used at the start and the end. Everything in between happens on grass.
I built it for people who like tinkering and wouldn't call it science. That covers kids, parents looking for an afternoon project, and teachers who want a reason to take a class outside. A paper plane is about the cheapest real experiment there is. It's also noisy enough to teach the habits that matter: change one variable at a time, write the prediction down first, repeat the test, and check whether your predictions were any good.
Demo
Paper Trials runs on your own machine (both models run locally), so there's no hosted demo. Setup is in the README: build the frontend, uv run papertrials, and optionally ollama pull gemma3:4b. Here's the loop in screenshots, using the built-in sample notebook.
The notebook suggests the next change. Each suggestion shows its predicted range, with your current best as a dashed line:
Test mode is the "go outside" screen. It shows exactly what to build, and the prediction is already locked in:
The Lab plots every prediction against what the plane really did:
Code
Surge77
/
paper-trials
A lab notebook for paper airplanes: plan one change, fly it outside, and let TabPFN and Gemma learn from every throw.
How I Built It
The backend is FastAPI and SQLite. The frontend is React with Recharts. Two open models do the interesting work, and plain code holds them together.
TabPFN predicts the distance
The hard part of this project is the amount of data. After a weekend you might have 40 flights. Most models can't learn anything from 40 rows, and the ones that can tend to be very confident and very wrong.
TabPFN is built for exactly this. It's a model pretrained on millions of synthetic tables, and it predicts from a small table of your own data in one forward pass, without any training. Each flight is a row: the plane's design, how it was thrown, and the wind. The target is the distance. I ask it for the 10th, 50th and 90th percentiles, and those become the predicted range.
The range is the important bit. Paper planes are noisy, so "14.3 m" would be a lie, while "probably 12.9 to 15.5 m" is something you can actually check. The Lab keeps score of how often the real median lands inside the predicted range.
Until there are 4 planes and 8 flights, TabPFN doesn't run. A nearest-neighbour guess fills in, its range widens the further a new plane is from anything you've flown, and the app labels it as a rough guess.
A gotcha worth knowing: newer TabPFN versions need a browser licence login before their weights download. My first version wrapped TabPFN in a try/except that fell back to the baseline, and it worked "fine". Then I noticed every single prediction in my sample data said baseline. TabPFN had never run once. The v2 weights download without an account, so the app pins those. After the fix, the sample notebook's TabPFN predictions landed inside their range 6 times out of 8, against 2 out of 3 for the baseline, and they missed by 0.8 m on average instead of 1.5 m.
The next experiment comes from code, not a chatbot
I didn't want a language model inventing plane designs. The recommender is under 200 lines of plain Python:
- Start from your best plane.
- Try every single-step change to one setting: a paper clip more or less, 1 cm wider or narrower, 5° sharper or blunter. Skip anything you've already built.
- Ask TabPFN about all of them at once and score each by predicted median, plus a bonus for a wide range. A plane the model is unsure about teaches it the most, so it's worth a flight even when it isn't the favourite.
- If your notes keep mentioning a stall, changes known to fix stalls (a clip on the nose, less tail bend) get a bonus too.
Gemma reads your notes, and gets checked
Gemma 3 4B runs locally through Ollama and has three small jobs.
Reading flight notes. "Went up, stalled and dropped" becomes the tag stall, which the recommender can act on. Gemma returns JSON constrained to a fixed list of issues. When Ollama isn't running, keyword matching takes over.
Writing a two-sentence lab note after each experiment. This turned out to be the most interesting part, because a 4B model writing about numbers will get them wrong in creative ways. Real examples from my testing:
"The median distance increased from 14.3 m to 14.1 m."
"The flight notes indicated a stall and veer right, which caused the median distance to fall below the predicted range."
"No changes were made from the previous plane, resulting in a veer left."
Every number in those is correct, so a "did it invent numbers?" check passes all three. So the code checks more than numbers. A note is rejected if it uses a number that isn't in the facts, says the distance went up when it went down, contradicts the inside/outside verdict, mentions a problem nobody reported, or claims a cause, because three throws can't tell you why something happened. A rejected note is replaced by a plain template. Each of those checks exists because Gemma did that exact thing at least once, and each has a test with the real sentence that triggered it.
Checking a photo. Before you throw, you can show Gemma a photo and it estimates whether the wings look even and the nose looks sharp. It's labelled as an estimate with a confidence score, it's never used by the prediction model, and the photo isn't saved.
Honesty about the sample data
The app ships with a sample notebook so you can look around before folding anything. Those flights are simulated: a made-up formula plus noise. The predictions inside them are real, though. A script replays the experiments in order through the app's own recommender and models, so every sample prediction was made before its flights existed. The sample is labelled everywhere it appears and one button removes it.
Why Does Open Innovation Matter?
It runs on a laptop with nothing to sign up for. There are no API keys, no account and no bill. A teacher can install it on a school laptop and a class can use it all term for free. Gemma runs on my GTX 1650 through Ollama, and TabPFN runs on the CPU, which is plenty at this data size.
The right tool for the job was an open one. Asking a hosted chatbot "how far will this plane fly?" gets you a confident sentence. TabPFN gets you a calibrated range from 40 rows of your own data. Because the model sits inside my code, I could ask it for quantiles, batch every candidate into one call, and keep score of its accuracy.
I could see and fix what went wrong. The licence fallback, the lab note that got the direction backwards: I found both because I could run the models locally as often as I liked and read every output. A rate-limited API makes you test less.
Your data stays with you. Your flight log, your notes and photos of your kid's paper planes stay on your own computer.
It's swappable. The Gemma model is one environment variable. If a better small model comes out next month, change one line and the guard checks still apply.
Prize Categories
- Best Use of TabPFN: TabPFN makes every prediction, with quantile ranges, from a dataset that starts at zero rows, and drives the experiment recommender.
- Best Use of Gemma: Gemma 3 4B runs locally for flight-note tagging, grounded lab notes and photo checks, all with tested guards against the mistakes it actually made.
Built with PriorLabs-TabPFN.



Top comments (1)
Locking the prediction before the three throws is a useful design choice. One statistical detail I would make explicit is the target of the 10th-to-90th percentile interval: an individual throw and the median of three throws have different uncertainty. Scoring a per-flight interval against experiment medians can make coverage look stronger than it is.
I would keep the three measurements grouped by experiment and evaluate whichever target the model actually predicts. Session-level grouping also helps separate changes in plane geometry from a different day's wind or throwing technique, while preserving the small-data workflow.