DEV Community

EazyHood
EazyHood

Posted on Fully Autonomous

SweepSheet: turn cleanup voice notes into reviewed counts

Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass Submission 🌿

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass.

What I Built

A cleanup has two different jobs: picking things up and keeping track of what was collected. SweepSheet puts the bookkeeping after the outdoor part.

Use your usual voice recorder to leave a short note: “Three plastic bottles and two cans.” Put the phone away. When you’re back, import the recording into SweepSheet. An open speech model transcribes it on your device, and a small parser proposes quantities in six categories. Listen, correct anything necessary, and confirm. Only then do the numbers enter your notebook and CSV.

The person I designed for is someone keeping the tally for a small cleanup, with several voice memos to turn into a table. This is a working prototype for that task, not a report of a cleanup I have already conducted. The examples are synthetic speech and are labeled that way throughout.

The most revealing test was a mistake. The sentence “Twenty three wrappers and twelve cans” came back as “23 rappers and 12 kids.” SweepSheet did not try to rescue it with another model. It left the counts for a person to enter after listening.

That decision shapes the whole application: speech recognition can help with transcription, but it does not get to approve the record.

Demo

Open SweepSheet — no account or API key needed.

SweepSheet in Edge after real local transcription and confirmation of the synthetic canal note: three plastic bottles, two cans, total five.

A real browser run with a clearly marked synthetic recording. The total is five because one note has been confirmed; the other two are still pending.

Choose Try the example notebook, then open Along the canal. The initial example transcript comes from the saved model test. Transcribe again runs Whisper on your own device; the first use downloads the model.

  1. Play the note and check 3 plastic bottles + 2 cans. Confirm: the total becomes 5.
  2. Review By the footbridge: 4 wrappers + 6 cigarette butts + 1 glass bottle. Confirm: 16 in total.
  3. Open An uncertain count. Its transcript says “Maybe 4 or 5 rappers.” The app holds the proposal; it adds nothing.
  4. Correct a confirmed quantity and explain why. The total changes by replacement, while the previous quantities remain in the review history.
  5. Download the confirmed CSV and the JSON review history. The latter includes the audio file’s SHA-256 identifier, original model output and corrections, but not the audio bytes.

The uncertain synthetic recording is held for review, with zero suggested counts and the previously confirmed total unchanged.

Uncertainty stays visible. Opening this note does not change the confirmed total.

Your own short audio files work too. Keep them below 60 seconds and 10 MB. A printable field card lists the six categories and a short recording example. There is no requirement to keep this app open outdoors or rely on a web recorder with a locked phone.

Code

Source, setup and evidence on GitHub · All ten test inputs and observed outputs

The app and repository were started on October 7 for this challenge. The application code is MIT licensed. Model and library licenses are credited separately in the README.

How I Built It

The pipeline is small enough to inspect:

Audio file → browser decoding → Whisper tiny.en in a worker
           → editable transcript → conservative count proposal
           → explicit human confirmation → totals and CSV
Enter fullscreen mode Exit fullscreen mode

I used Transformers.js 3.8.1 and the quantized Xenova/whisper-tiny.en ONNX conversion, pinned to revision 79fb389fc764e7c395bd330e9531d9d32ada7049. The browser runs inference through ONNX Runtime WASM. A worker keeps that computation away from the interface. The app stores recordings, transcripts and review history in IndexedDB; there is no transcription server or application backend.

The parser supports exact quantities in six categories. It holds corrections, vague amounts, negation and unsupported numbered categories. “Bottles” alone needs a material. The parser does not silently repair “rappers” to “wrappers.” These are inspectable, intentionally limited rules, not a claim to understand every sentence.

Every imported note starts pending. Exact-byte reimports are detected with SHA-256, so importing the same file twice does not double the tally. Corrections replace a note’s quantities instead of adding them again. Excluded notes remain available for reference. The CSV contains confirmed records only; the JSON also retains pending and excluded notes. Spreadsheet-sensitive text is escaped.

What the tests actually showed

I generated eight short spoken fixtures locally with Microsoft Zira Desktop, plus digital silence and deterministic white noise. The full method distinguishes synthetic speech, CPU inference and browser execution.

Test group Observed result
Four unambiguous count sentences Three produced the intended quantities; one was held after a recognition failure
Four correction, uncertain, vague or not-collected sentences All held for manual review
Silence and white noise Neither proposed counts; noise was transcribed as “(water splashing)”

The noise result matters too: even a speech model can invent a plausible description. Nothing about that output establishes that there was water, much less collected litter.

The tests also check that pending and excluded notes never affect totals, corrections retain the earlier values, invalid quantities are rejected, and CSV text cannot become a spreadsheet formula. These checks protect the recording workflow; they do not establish outdoor ASR accuracy.

The current limits are practical. It is English-only. Browser audio-format support varies. Model preparation needs internet, and offline page reload is not supported in this version. Browser storage can be cleared, so keep original recordings and export your work. Wind, real voices, accents and mobile performance still need field testing. I have not measured a reduction in screen time or claimed environmental impact from these example counts.

AI assistance: Codex substantially assisted with implementation, tests, documentation and this write-up. The synthetic demo voice is disclosed separately. No field experience, interviews or user metrics have been invented to fill in the story.

Why Does Open Innovation Matter?

The useful open component here is the transcription model, not a logo beside a text field. Without it, the user must type the recording into the notebook. With it, audio can be transcribed on the same device, without uploading cleanup recordings or obtaining an inference API key.

The pinned weights and readable pipeline also make the failure above reproducible. Someone can try another open speech model against the same WAV files, or adapt the parser to another cleanup vocabulary, while keeping the confirmation rules intact. The application does not depend on a remote provider continuing to offer a particular model or pricing tier.

A closed API could also transcribe these notes. The reason to use open weights is control over where inference happens and the ability to inspect and replace the component. It does not make recognition magically reliable.

Cleanup logging itself already exists—Clean Swell is one example—and a voice memo plus a spreadsheet is a reasonable manual alternative. SweepSheet’s contribution is narrower: keep the recording beside a proposed tally, make uncertainty visible, and preserve what changed after review. The next useful test is taking that complete workflow to a real cleanup and finding out whether the bookkeeping is actually easier.

Top comments (1)

Collapse
 
launchgatecheck profile image
Launch Gate •

The replacement rule is worth testing at the storage boundary too. What happens if the same confirmed note is open in two tabs, and each tab saves a different correction from the same starting revision?

A fixture could confirm 3 bottles + 2 cans, then save corrections to 4 and 6 bottles from those two tabs. A stale save should either be rejected for review or have an explicit conflict rule; the tally and review history should agree on the accepted quantities after a reload. That would test a different failure from exact-byte reimports. I haven't run SweepSheet; this is a suggested IndexedDB/review-history fixture, not a claim that it currently loses edits.