This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass
What I Built
A doorstep, a park bench, a patch of shade. Pick somewhere ordinary and give it a minute. What comes into focus?
SoundWalk turns five seconds of sound into a reason to listen for sixty. Record your surroundings, let a local AI model suggest a listening direction, and choose what to follow. At the end, bring back one sentence in your own words.
I built it around a handoff. The model supplies the starting point. The listener takes over.
For water, the invitation is: “Find one small sound inside the flow.” For street sounds, it asks you to follow something as it approaches and fades, then hear what emerges in the gap afterward. A familiar place becomes a small listening exercise.
Once you start your minute, the page dims. The microphone has already stopped. Read the invitation, look up, and let the place supply the experience. A soft chime brings you back.
Your observation becomes a Sound Card: a little keepsake with the model's suggestion, your chosen direction, and the words you wrote. That is the ending I wanted—a specific detail worth taking home.
Demo
Take a soundwalk → · Watch the 77-second browser walkthrough →
- Prepare the local model while you have a connection. The first visit downloads about 18 MB.
- Find a place outside to pause, then record five seconds.
- Choose a listening direction and start your minute. Keep the tab open while the screen dims; changing tabs pauses the timer.
- Write one observation. Save your card in the browser, export a PNG, or share the words.
At your desk, Try a shoreline sample offers a complete rehearsal using a credited sea-wave recording. The walkthrough follows this route with actual local inference, an edited view of the sixty-second timer, and a clearly labeled scripted observation.
Sample rehearsal. The card's observation is a scripted example; the credited five-second sea-wave clip loops during the listening minute.
If you try it outdoors, share your phone/browser, the direction you chose, and one thing you noticed. That is the feedback I most want to build on.
Code
Source, local setup, and credits →
SoundWalk was designed and implemented with Codex, OpenAI's coding agent. The application and listening invitations are MIT-licensed. YAMNet and TensorFlow.js are Apache 2.0; their files and attribution ship with the repository.
How I Built It
Turn a classification into a question
YAMNet scores 521 audio-event classes. SoundWalk maps relevant labels into a smaller vocabulary: birds, water, wind, street rhythms, and life nearby. Each group uses its strongest matching label, keeping related labels together.
Then comes the important decision: the person chooses the direction. A bus can dominate five seconds of audio while a quieter rustle holds your interest. Choosing the rustle is already an act of attention. The final card preserves the model's suggestion and your choice as separate fields.
Weak or silence-dominant matches lead to an open invitation: find your nearest sound, find your farthest sound, and move your attention between them.
This is the design work around the AI. “Water” gives the activity context. “Find one small sound inside the flow” gives the listener something to do. The prompts are authored, short, and visible in audio.mjs.
Pack the model before heading outside
Five-second microphone capture
→ mono audio at 16 kHz
→ YAMNet in a browser worker
→ listening suggestions → your choice
→ sixty seconds → your observation
An AudioWorklet captures the audio; the input track stops after five seconds. Resampling prepares it for YAMNet. TensorFlow.js runs inference on the CPU in a Web Worker, keeping the interface responsive.
The model files total about 16.14 MB. Making that first download visible became part of the experience: a preparation screen before the walk.
A service worker caches the app, runtime, model graph, four weight shards, labels, and sample. When the footer reads “Ready for an offline soundwalk,” that browser can reopen the experience offline. Audio processing happens on the device. The latest fifty cards live in browser storage; exported images are ordinary files you can keep.
Recording the demo caught a useful browser wrinkle: the cached WAV decoded for inference, while native media playback rejected its URL. Creating a local Blob from the same bytes fixed playback, including after an offline reload. A playback failure now pauses the rehearsal and offers a Resume action. Hearing the sample is part of making the invitation work.
Give development a repeatable shoreline
The bundled sample made the whole sequence reproducible. In Chromium 153 on Linux, its first YAMNet inference took 4.23 seconds on the CPU and led to the water invitation. The rehearsal completed sixty seconds, saved a card across a reload, and exported a PNG.
An offline run exercised the microphone path with Chromium's file-backed audio simulation. Selecting Wind after the model suggested Water checked the handoff I cared about: the saved card retained both. The setup and results are documented here.
The sample comes from HerbertBoland's “Branding_kort.wav”, via ESC-10, under CC BY 3.0. Its source label follows it through the rehearsal and saved card.
Why Does Open Innovation Matter?
Open weights made the product's geography possible: the model travels with you. Once prepared, SoundWalk can do its work on the listener's device, wherever that browser is taken.
Open code makes the next layer just as accessible. You can trace an audio label into a category, read the invitation it suggests, and change the experience in a small edit. A school-garden version could keep the same model and rewrite the invitations around its own sounds and language.
That is the kind of contribution I would love to see: someone who knows a place writing a better question to ask there.
The open model supplies a shared starting point. Every walk ends with something particular: one person's attention, in one place, for one minute.
Prize Categories
Best Use of GitHub Copilot — GitHub Actions automation. The category explicitly includes “automate your project with GitHub Actions.” SoundWalk's workflow runs audio tests, checks licensing provenance and required model assets, and verifies the offline dependency pack before deploying to Pages.
That gate protects the experience described above: a missing weight shard stops publication. The release verification and deployment is public.

Top comments (1)
tr.ee/dev-to
Some comments have been hidden by the post's author - find out more