This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
Somewhere on Sam's laptop, there's a riff Sam swears is the best thing he has written all year.
It's in a file called recording_027_final_v2.wav. Or maybe heavyyy_enough.wav. Or idea_actually_good.m4a. He has 100s of these, and not many filenames say what's inside.
Here's the cruel part: you don't remember a riff by its name. You remember how it felt. Ask Sam about the lost one and you get something like:
"heavy, Drop C, around 145, I think it was a chorus idea"
That's a perfectly good search query. It's also one that no folder on earth can answer.
So I built something that can.
RiffSalad (yes, that's what the recordings folder looks like) is a local AI vault for guitarists. You drop riffs in, or record them straight from the browser, and it listens: tempo, key, energy, how busy the playing is, even a MIDI transcription. It writes a short description and a few tags for each one. Then you search the way you actually think:
heavy drop C riff around 140 BPMclean mellow fingerpicked thing in E minorsomething like my last blues riff
You can talk to your riffs, too. Record a voice note ("this was the bridge, needs a key change") and it's transcribed on your own machine and becomes searchable. There's also a waveform player, inline editing, and a "find similar riffs" button.
Who I built it for
Sam is an intermediate guitarist whose head is constantly popping off with new ideas. The last time I saw him, he was sitting in front of his laptop, humming a tune at it, while I sat there being the overachieving automator. I wanted to build the thing that lets him stop scrolling and just type what he remembers. Mostly because I wanted him to stop humming. lol!
What Sam said
Finallyyy, I can have something organized in my life and easy to find.
Demo
Code
Your guitar ideas, searchable.
A local-first AI vault for guitarists. Record a riff, import the file β and find it later by describing it in plain English.
The Problem
Every guitarist knows this feeling: you recorded a cool riff a few months ago, but you can't find it because you saved it as recording_027_final_v2.wav. You remember it was a heavy, slow, drop-tuned thing around 80 BPM β but searching a folder of audio files for that is impossible.
RiffSalad solves this. Import your recordings and search them like you'd describe them to a bandmate.
"heavy drop C riffs around 140 BPM"
"that mellow fingerpicking thing in E minor"
"fast aggressive shredding, lots of notes"
How It Works
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Import WAV / MP3 / M4A β
β β β
β βΌ β
β librosa analyzes: BPM, Key, Energy, Note Density, Duration β
β β β
β ββββΊ Ollama LLMβ¦Everything runs locally and needs no API keys:
git clone https://github.com/jesi318/Riff-Salad
cd Riff-Salad
cp .env.example .env
docker compose up
Then open localhost:3000. The first run pulls the models. After that, it works with no internet.
How I Built It
One decision shaped everything: the small model does as little as possible.
I wanted this to run on a normal machine, so the language model is a 1.5B-parameter qwen2.5 served by Ollama. That's tiny. A model that size is good at turning "heavy drop C riff around 140" into structured filters, and bad at deciding whether a riff actually is heavy. So I stopped asking it to decide.
The rest of the stack is deliberately plain:
-
Ears:
librosameasures BPM (harmonic/percussive separation, then tempo tracking), key (matching against tonal profiles), energy, and note density. A Basic Pitch-style ONNX model extracts MIDI. -
Voice:
faster-whisper(thetinymodel) transcribes voice notes locally. -
Memory:
nomic-embed-textembeddings via Ollama, stored alongside a SQLite database. - Plumbing: FastAPI and SQLAlchemy on the back end; React 19, TypeScript, Vite, Tailwind, and WaveSurfer.js on the front.
- On an M4 mac air, a riff takes less than 15 seconds to analyze.
Search runs in four stages, and the chat model only gets the first one:
- Parse. The LLM turns your sentence into structured filters: BPM range, key, energy, mood words.
- Prune. Plain code throws out anything that breaks the measurable facts. "Around 100 BPM" still lets 96 through.
- Re-rank. Embedding similarity sorts what's left.
- Fall back. If the LLM or embeddings are unavailable, it still matches text across titles, notes, tags, transcripts, and descriptions.
The rule I'm proudest of: the tagger isn't allowed to say "heavy," "metal," or "doom" unless the numbers back it up. Small models love to over-claim, and "heavy" is the easiest word in music to overuse. So the measured facts get the first vote, and the model just gets to phrase things nicely. "Find similar" follows the same philosophy, blending embedding similarity with plain feature distance (BPM, key, energy), so "sounds like this" is something you can sanity-check.
It isn't magic. Key and BPM detection are heuristics, and sometimes they're wrong, which is why both can be overridden inline.
Why Does Open Innovation Matter?
A riff is a diary entry you can hum. It's half-finished, sometimes embarrassing, often unreleased. I didn't want Sam to upload his unfinished ideas to someone else's server just to find them again later. With open models running locally, they don't have to: the audio, the transcripts, and the database all live in a data/ folder on his own machine.
Open also gave me things a closed API wouldn't have:
- No meter running. Every import triggers analysis, tagging, and embeddings. Sam can throw 200 riffs at it without a bill, a quota, or an API key to babysit. After the first run, it works in a rehearsal room with no wifi.
-
Models are a config line.
LLM_MODEL,EMBEDDING_MODEL, andWHISPER_MODELare environment variables, so swapping a model never means touching code. - Constraints made the design better. Because I committed to a tiny local model, I had to build the hybrid pipeline: the LLM for language, plain code for facts. That's why search still works when the model is down, and why I can explain every result. A giant hosted model would have let me be lazy, and I suspect the app would be less predictable for it.
Sam still has 103 files called recording-something. He just doesn't need the filenames anymore.
Top comments (0)