DEV Community

Cover image for Nimbus Choir: Point Your Phone at the Sky, Hear It Sing
Ch Yasaswini
Ch Yasaswini

Posted on

Nimbus Choir: Point Your Phone at the Sky, Hear It Sing

Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass Submission 🌿

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass

What I Built

Nimbus Choir points your phone's camera at the sky and turns it into a soft, generative ambient piece -- entirely offline, no account, no cloud service watching what's above you.

Every cloud the camera tracks becomes one voice in the choir:

  • Cloud type picks the timbre (cirrus, cumulus, stratus, cumulonimbus each hum differently)
  • Size picks volume
  • Height picks pitch
  • Left-right drift pans it across the stereo field
  • A short-term weather forecast tunes how consonant or unsettled the whole piece feels

Nimbus Choir's main screen -- camera stage, controls, and the

The intended use is exactly what the theme asks for: point it at the sky, tap once, then put the phone down. The screen is the shortest part -- after that it's just you, outside, listening to clouds drift while an ambient pad drifts with them.

There's also an optional Earth mode: tap the globe icon, watch a few cloud shapes part to reveal a real spinning globe (actual coastlines, actual rotation -- written from scratch, no mapping library), tap a glowing region, and the same sound engine sonifies live NASA satellite imagery instead of your camera. It's clearly marked "Online" and falls back to the camera automatically if the network's unreachable.

The globe picker -- real coastlines, orthographic projection, snaps to the nearest live-coverage region

Demo

Live: web-pi-wine-28.vercel.app -- the real frontend (camera pipeline, tracking, sound engine) running live. This is a static deploy of just /web, so the backend-dependent bits (tension forecast, sky narrator, Earth mode's satellite proxy) aren't reachable from it and gracefully fall back exactly as they're designed to offline -- the core camera-to-sound experience is fully live and functional. For the complete experience including those, clone the repo and run the FastAPI server alongside it (instructions below).

Code

github.com/Yasaswini-ch/Nimbus-Choir

How I Built It

/web -- vanilla JS + Vite. A real-time sky segmentation pipeline (sky masking, connected-component cloud detection, an IoU tracker that handles merges/splits as clouds drift and overlap) feeds a cloud-type classifier that runs entirely in-browser via onnxruntime-web -- a frozen, open-weight DINOv2 backbone with a small trained head, quantized to ONNX. Zero network calls for the core experience. The sound engine is hand-built Web Audio: detuned oscillator pads through a lowpass + slow tremolo, plus physically-modeled partial-based chimes.

/training -- the classifier pipeline. Public CCSN cloud dataset → dedupe/augment → DINOv2 (open-weight, frozen) + trainable head, distilled with optional Tinker teacher soft-labels for extra accuracy, but the version actually shipped here trained on ground truth only -- $0 spent, no closed model touched at all. (CCSN has no clear-sky class, so I generated 60 synthetic gradient-sky images as a documented stand-in for that one label.)

/forecast -- an open-weight TabPFN model (a tabular foundation model) predicts the "tension" value from cloud cover + pressure trend, trained on public Open-Meteo reanalysis data across ~40 global locations. Held-out-by-location accuracy (tested on 6 cities -- Buenos Aires, Cairo, Mumbai, São Paulo, Shanghai, Sydney -- the model never saw during training): 77.5%.

/server -- FastAPI. Serves the tension model, an optional sky narrator (local Gemma 3 via Ollama writes a caption, Piper -- open-weight, fully offline TTS -- speaks it), and Earth mode's NASA GIBS satellite tile proxy (confirmed live against GOES-East, GOES-West, and Himawari-8, no API key needed).

Why Does Open Innovation Matter?

This is the part I actually want to highlight, because it wasn't just philosophical -- it changed a real decision mid-build.

The narrator originally used ElevenLabs for voice. It worked, but it meant: no network → no voice, an API key to manage, and a per-character cost for something meant to run in a field with no signal. I swapped it for Piper, an open-weight local TTS. Same user-facing feature, zero network calls, zero marginal cost, and it keeps working exactly where the app is meant to be used -- outside, possibly with no bars.

That one swap is the whole argument in miniature:

  • Runs with no internet. The main camera mode -- segmentation, tracking, classification, the entire sound engine -- never makes a network call after first load.
  • Nothing about your sky leaves your device. No photos, no GPS (even the forecast lookup coarse-rounds coordinates to 0.5° before it ever leaves the client).
  • Costs nothing to run, at any scale. The open-weight classifier, TabPFN forecaster, and Piper narrator all run locally after their one-time download. No per-request bill waiting at the end of a demo day.
  • The one online feature is opt-in and honest about it. Earth mode needs the network (real satellite data has to come from somewhere), so it's the only part with an "Online" badge, and it's the only part that can't run in a field with no signal -- which is exactly why it's optional, not the default.

Open-weight models didn't just make this cheaper to build. They're the reason the app can keep its core promise -- point it at the actual sky, anywhere, with nothing phoning home -- instead of quietly becoming "offline, except for the parts that aren't."

Prize Categories

  • Best Use of TabPFN -- the /forecast tension model, trained and honestly validated as described above.
  • Best Use of Gemma -- the sky narrator's captions, generated by a local Gemma 3 model via Ollama.

Top comments (0)