DEV Community

Cover image for Open When: a letter that only opens when the world says so
Gowsi S M
Gowsi S M

Posted on

Open When: a letter that only opens when the world says so

Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass Submission 🌿

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass

What I Built

Open when you miss me. Open when you can't sleep. Open when it finally works out.

If someone has ever handed you one of those envelopes, you know the feeling: a small weight of somebody's care, waiting for the right moment.

But there's a catch. Nothing stops you from opening that envelope early, even on the bus.

Open When keeps the envelope and changes the lock.

OPEN WHEN YOU HEAR RUNNING WATER.

You write a message, choose a sound and seal it. You get a link. The person you send it to sees a sealed letter and that single line. They tap Begin, put the phone in their pocket and go for a walk.

While they walk, a small audio model on their phone listens to the world. When the sound is really there (a tap, a stream, a fountain), the screen turns green, the envelope opens and your words appear.

The idea is simple: you can't just open the letter at your desk. You have to take it outside and find the sound.

Two rules shaped everything:

  1. The AI never writes a word. The message is yours.

  2. The AI gets exactly one job: listen. It says whether the sound is present. It does not decide when the letter opens. A plain, readable piece of code does that and it refuses to be rushed.

Demo

Live demo: open-when-gp8x.onrender.com

The walkthrough:

  1. Create: write the message, pick a sound, pick a letter style.
  2. Seal: the envelope folds shut and a seal is pressed. You get a link.
  3. Open the link on a phone. The first time, it downloads the listening model (about 161 MB, once).
  4. Begin, then pocket the phone. The screen fades to a single dark dot.
  5. Listen and walk. Soft, subtle cues tell you if you are getting warmer.
  6. Confirmed: the dot turns green and the envelope opens.

Code

Open When

A letter that only opens when the world says so.
Sound-locked messages, decided by an audio model that runs on your phone. No server. Your audio never leaves the device

Live demo on Render React 18 TypeScript 5 Vite 5 CLAP audio model On-device inference Touch Grass

The idea · How it works · Run it · Privacy


The idea

Most "go outside" apps tell you to go outside. Open When gives you a reason: someone left you something, and the only key is a place.

  • The world is the key. There is no timer, button or password. The real environment is the trigger.
  • The AI never writes the message. It only listens and decides whether the sound is present.
  • The audio never leaves the device. Listening happens in short windows in memory and is discarded.
  • No server. The letter lives in the link itself (the URL hash). There is no database, account or backend.

How it works

mic -> 5 s audio windows ->
…

How I Built It

mic -> 5 s audio windows -> CLAP audio model (in the browser)
    -> score vs. rival sounds -> smoothing + "held for several seconds"
    -> confirmed -> the envelope opens
Enter fullscreen mode Exit fullscreen mode

The model. CLAP is an open model that puts sounds and sentences in the same space, so "running water" can be compared with what the microphone hears. I run an ONNX conversion (Xenova/clap-htsat-unfused, 8-bit quantized) with Transformers.js, entirely in the browser. There is no inference server.

The letter is the link. The message, the condition and the letter style are encoded into the URL hash. There is no database, no account and no backend. (The flip side: anyone with the link can read the message, so it is for people you would hand the paper envelope to anyway.)

Honest detection. A model being loaded does not mean it is calibrated. CLAP was happy to hear "running water" in one recording and miss it in another, so I stopped trusting a single score:

  • Several phrasings per sound, summed: flowing water, a stream, water over rocks, a fountain, a tap.
  • Rival sounds (wind, rain, traffic, music, silence) are scored alongside, and the target has to beat the best rival by a clear margin.
  • The unlock has to hold. One lucky window is not enough; the signal must be strong across several windows and sustained for a few seconds, with a cooldown after it breaks.
  • A state machine decides, not the model. The AI hands over evidence; deterministic code judges it. I can unit-test that code without a microphone or a model.

A debug panel for real life. Open the app with ?debug and it shows the mic state, signal level, each sound's score, the margin over the rivals and the detector's state. You can copy a CSV log of a walk, and score a recorded file through the same pipeline, which tells you whether a failure is the microphone or the model. This is how I stopped guessing and started measuring.

The stack: React 18, TypeScript, Vite, Transformers.js, a service worker so it installs like an app, and static hosting on Render.

Taking it outside

I tried Open When with different sounds, and running water was detected well enough to trigger a letter. People talking was trickier, though. The model struggled to recognise it consistently, which reminded me that detecting sounds in the real world isn't always as straightforward as it seems.

The listening process runs locally in the browser after the model loads, so the sound analysis doesn't need a server handling every recording.

Why Open-Source AI?

Think about what this app listens to: the sounds around a person who is walking alone, often somewhere private. A cloud audio API would mean streaming that to a server I don't control for the whole walk.

With an open-weight model, the whole loop stays in the phone:

  • Your audio stays yours. It is analysed in short windows on the device and isn't uploaded to a server.
  • I can open the box. When CLAP missed a sound, I could change the phrasings, add rival sounds and inspect the actual scores.
  • It costs nothing to run. No per-minute audio bill, no API key, and the whole app is static files.
  • The model is swappable. The rest of the app only talks to a small matcher interface, so a better audio model can replace CLAP without touching the letter, the seal or the unlock.

The honest trade-off: CLAP is general-purpose and not perfect. A bigger hosted model might hear faint water more reliably. I chose a smaller model that I can trust with a stranger's walk, then made the detection around it strict instead of hopeful.

What I Learned

The hardest part was not the model. It was deciding when the model is allowed to be believed.

An app like this can fail in two bad ways: the letter never opens (frustrating) or it opens at the desk (the whole point is lost). Most of the work went into making the second one hard and the first one rare, and into being honest about which one I was looking at.

And one small thing I did not expect: the best screen in the whole app is the one where you stop looking at it.

Prize Categories

  • Best Use of Render: Open When is hosted as a static site on Render. HTTPS lets the microphone work, and the sealed-letter link opens directly on a phone. Render makes the experience accessible without needing a separate backend.

Maybe the nicest thing about a letter is that it asks you to slow down a little.

If you could write an “Open When” letter for someone in your life, what would you want them to hear and what sound would you ask them to find before they could read it? 💌

Top comments (0)