This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass
What I Built
My reading system is a row of browser tabs I opened because the title looked worth my time, then kept open for weeks, because reading takes a chair and a chair is where I spend the day already.
Tabs to Trails turns those tabs into walks. A bookmarklet called "Walk this tab" drops the page you're on into a list, and each row shows how many minutes it takes to hear. When I've got twenty minutes, I pick 20. The app ticks what fits, oldest first, says how each piece will be read ("in full", "condensed from 27 min", "part 1 of 3"), and my laptop makes one MP3 of about that length. Gemma 4 rewrites the text for the ear and the Kokoro voice reads it, both on my machine. I scan a QR code with my phone, download the file, and put the phone in my pocket.
Everything the walk needs is inside the MP3: a short intro, the piece, a chime at the measured middle with "You're halfway. If you're walking out and back, turn around now.", and one question for the way home. The phone needs no app and no Wi-Fi once the file is on it. I set one rule for the design: the app has to be done with me before I reach the door. The only dark screen in the product is the phone's player, and it says "Pocket your phone." and little else.
Three things set it apart from "article to podcast". You bring the content, the thing you already meant to read. The length of your walk shapes it. And it stays faithful: the voice reads prose that fits as written, Gemma condenses only what doesn't fit, and a number guard checks every number in the rewrite against the source. Code blocks and tables are described rather than read, because a code block read aloud is noise.
Demo
Listen to two finished walks on the project page: my own DEV post, asked for 10:00 and measured 10:15, and part one of Thoreau's "Walking", asked for 20:00 and measured 19:39.
A note on the article in the video. It is "Every Software Developer Has Blamed…" by @sylwia-lask. I picked it as a real-world test, and about 20 seconds of the audio version made it into the video. I hope Sylwia doesn't mind me using her post as the example 😇😇😇
Go read the full post:
Code
nazboyko
/
tabs-to-trails
Turn your reading backlog into a walk: a local Gemma 4 and Kokoro make an MP3 sized to your walk.
Tabs to Trails
Turn your reading backlog into a walk.
Save tabs now. Walk them later.
Tabs to Trails keeps what you mean to read, measured in walking minutes. When you have twenty minutes, pick 20 and it proposes what fits. Your computer turns that into one MP3 of that length: Gemma 4 rewrites the text for listening and the Kokoro voice reads it, both on your own machine. You scan a QR code, put the phone in your pocket and go.
The file holds everything the walk needs: a short intro, the reading, a chime and a cue at the halfway point so you know when to turn around, one question for the way home, and a sign-off. Its chapters let the phone's own player resume and skip, and an audiobook copy (.m4b) is one tap away. Once the file is on the phone, the walk needs nothing else.
…
Node and TypeScript: a Hono server, a React front end, Gemma 4 through Ollama, Kokoro-82M through kokoro-js, ffmpeg for the MP3. Gemma 4 and Kokoro-82M are Apache-2.0; the full credits with licenses are in the README.
How I Built It
The pipeline:
link or text, saved to the list and read once
-> picked by the walk's length
-> plan: in full, or condensed to a word budget
-> Gemma 4 rewrite per section, number guard
-> Kokoro voice per sentence group
-> measure, fit once, place the halfway cue
-> MP3 + script + QR code
Every stage writes into one folder per walk, so a build can stop and pick up where it was. I tested that with kill -9 in the middle of the voice stage; on restart the server said "Picking up 1 unfinished walk" and carried on from the third section.
Twenty minutes is not twenty minutes
The length is the product, so I measure it instead of trusting a words-per-minute guess. That guess was wrong in two places.
The voice first. My calibration read a smooth 250-word passage at 188.7 words per minute, and then my own DEV post at 10 minutes came out at 10:42. Across that one post the same voice went from 130 to 174 words per minute per section, because numbers, colons and short paragraphs cost time, while characters per second stayed between 14.3 and 15.9. So the app now calibrates in characters per second, on a passage with numbers in it, and converts that into words per minute for each source.
Then the model. Gemma does not hold a word count on a mild cut: asked for 697 words of a 939-word section, it returned 1,027, longer than the source. Short targets went the other way: asked for 823 of 1,037 words of Thoreau, it wrote 450. So each section gets one guarded shorten or lengthen pass, and whatever it over- or undershoots carries into the next ones. At the end, a fit pass looks at the measured audio once: more than 5% over, it shortens the largest condensed section; more than 8% under, it gives time back to the section that left out the most.
The sample above took my DEV post from 1,726 words to 1,531, with 4.1 seconds of model time and 63.5 seconds of voice on an Apple M5 Max, for 10:15 of audio. Kokoro reads about nine times faster than real time here, so most of the wait is the voice. The app says "Lace up." while it works, and means it.
The number guard
The prompt tells Gemma to use only what is in the source. A small model adds a number from time to time. So after each rewrite, code compares every number in the output with the numbers in that section of the source, spelled-out ones included. A new number sends the section back once, naming it. If it comes back again, the script view draws a dashed outline around it: "Check this number."
The guard once flagged a number that was right: Wikipedia's "one-hundred and fifty percent" came back as "150 percent", and the guard read it as two numbers. It adds them up now. It also misses what it cannot see: "the kid screen got the time" (my evening, in other words) came back as "the kid screen showed the time", with no number involved. The script view's footer says to review anything important, and I mean that.
Read-along that lands on the word
Kokoro speaks in groups of a few sentences, and the app knows where each group starts. Inside a group, my first version spread the time over the sentences by their characters. Measured against the pauses in the audio, those starts were a median 0.3 seconds off, enough to mark the wrong sentence. Now each start moves to the nearest real pause in the voice, preferring longer ones. In the two samples, 231 of 232 sentence starts sit on a pause of at least a quarter second, 40 milliseconds before the voice is loud again.
What broke
- Kokoro reads at most 510 phoneme tokens per call and drops the rest without a word. A 294-character chunk of money figures was 1,237 phonemes. Each chunk is now measured with the same phonemizer and split until it fits.
- Wikipedia read its footnotes aloud: citation marks like [12] became "twelve". The app strips reference lists and footnotes now, and names what it left out.
- The first run read an HTML comment. My own post ends with
<!-- Thanks for participating! -->, and the voice said it.
I took it outside
I didn't film the walks; the video above is a screen recording with the real audio. The 10- and 20-minute walks went without a hitch. On a longer one, a phone call interrupted "Play it here" and I had to open the page and start again, so for long walks the downloaded MP3 or the audiobook file is the better way.
What I cut and what's next
I cut a model comparison with gemma4:e2b, because it wasn't on my laptop and I didn't want to download a model for one table. Next I'd try a pace setting per voice, and test the audiobook file and the reminder on more phones than mine.
Why Does Open Innovation Matter?
The documents I most want to walk with are the ones I can't put in a demo. With open weights on my own machine, the text goes from the clipboard to my disk and nowhere else, so the app works for anything I'm allowed to process on my own machine.
It also changes what a walk costs. There is no meter, so I can make one for a ten-minute errand and throw it away. Paste mode works with the network off once Ollama and the voice are installed.
And anyone can change it. The model is one setting (OLLAMA_MODEL), the prompts are plain text in the repository, and the voice is an Apache-2.0 model you can swap.
Prize Categories
Best Use of Gemma. Gemma 4 (gemma4:e4b through Ollama) does the language work: it condenses sections to a word target, describes code and tables for the ear, rates how much each section matters, and writes the closing question. The code around it keeps it honest: word budgets, the number guard, and a measured fit to the walk's length.




Top comments (2)
Hahahaha, I love this! 🤣 Now I know what I'll be doing on my walks! And thanks for the mention, I'm officially famous now! 🤣
Famous and now available in audio! 🎧 Your post came out as a 7-minute walk, which is exactly one lap around my block. 😇 Thanks for being such a good sport about it, and for writing something worth walking with.