DEV Community

Cover image for Listening Loop: three stops, six seconds, and an open model that stays on your phone
Kanishq Sharma
Kanishq Sharma

Posted on Fully Autonomous

Listening Loop: three stops, six seconds, and an open model that stays on your phone

Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass Submission 🌿

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass.

What I Built

Listening Loop is a field notebook for your ears: choose three places to pause, record six seconds at each, and write down what you actually noticed. An open audio model suggests sounds, then the app asks you to put the phone away and walk somewhere different.

The design question was: how little screen time does a useful AI walking app need? I chose a short loop with no map, score, streak or infinite feed. A doorway, a green corner and the way home are enough. You can adapt those stops to your neighbourhood and mobility.

The model can help you notice a sound you would otherwise pass by. It cannot tell you whether a route is safe, identify a bird species or replace your own observations. The journal deliberately asks: What did you actually notice?

Demo

Open Listening Loop

Press Prepare for a walk while online. Wait for the offline-readiness message; then the app and its model can run without a connection. The initial download is about 19 MB.

To try it at your desk, select Analyse example birdsong. This uses Oona Räisänen's public-domain chaffinch recording from Finland, not a recording I made outdoors. It never adds a stop to your journal.

For a walk, name your first stop and press Listen for 6 seconds. Microphone permission is requested at that point. Add a note, move somewhere different, and repeat twice. The finished journal downloads as JSON. Remembering it on the device is optional.

Imported audio is supported for testing and accessibility, but every imported stop is labelled as imported, not a verified outdoor observation.

Code

Public repository

Original application code and SVG artwork are MIT licensed. Google YAMNet and TensorFlow.js retain Apache 2.0 licences. Source links, notices and model-file hashes are included. This project and repository were created on October 6, during the Week 1 window.

How I Built It

The core is Google YAMNet, an open audio-event model with 521 labels, running through TensorFlow.js. The app averages stereo imports to mono, resamples to 16 kHz and averages the model's frame scores across a short clip. Broad Nature, People and City summaries use editable label groups.

A label is a suggestion, not a certainty. YAMNet scores are not calibrated probabilities. Quiet audio, clipping and ambiguous results receive different prompts instead of a confident-sounding claim.

The supplied birdsong gave a Bird score of approximately 0.556 and a Nature summary. That number is a model score, not a 55.6% confidence estimate. Silence produced a Silence label and a warning that the moment was too faint.

All model weights and the browser runtime are served as static files. A service worker caches a fixed list after successful preparation. There is no inference server, location request, analytics script or audio upload endpoint. Capture tracks are stopped after recording or cancellation; only notes, sources and model suggestions enter the journal.

Validation includes eight unit tests and twelve Chromium integration checks. The browser checks exercise the actual YAMNet model, permission denial, recording cancellation, a complete generated recording, invalid imports, silence, journal export, storage opt-in and forgetting, offline reload and a 390px layout. They also check for browser errors and network-upload requests.

Two details mattered during testing: a cancelled capture must release its track, and an example result must never quietly become a claimed walk observation. The integration tests check both. Repeated inference also returned to the same persistent tensor count.

Limits: the microphone test uses a generated MediaStream. I have not completed an outdoor, physical-phone or Safari trial, so I am not claiming one. Browser storage can be evicted, and overlapping outdoor sounds may confuse the model. Those are the next practical checks.

Why Does Open Innovation Matter?

A walk should keep working when the signal disappears. Having distributable model weights made offline inference possible without an API key or subscription. It also means a person's six seconds of sound do not need to leave their device.

The open implementation lets someone inspect the input processing, change the sound groups, adjust uncertainty thresholds or replace the model. They can run this static app on their own host. The tradeoff is a larger first download and work constrained by the device; that is visible in the preparation step rather than hidden behind a server call.

Build Process and Attribution

The implementation, documentation, tests and this write-up were generated by Codex under my instructions. DEV's AI disclosure is set to Fully Autonomous. I did not train YAMNet or collect the example recording. The test results are software evidence, not a claim of a human field trial.

Model: Google YAMNet. Runtime: TensorFlow.js. Example audio: Oona Räisänen, Fringilla coelebs, 2007. Full notices are in the repository.

For this entry I am entering the overall challenge; I am not claiming a partner-technology category.

Top comments (0)