DEV Community

PartFit 3D
PartFit 3D

Posted on

Running Spotify's basic-pitch in the browser for audio-to-MIDI

I wanted a quick MIDI sketch of a melody I had sung or played, without uploading the recording to anyone's server. Spotify's basic-pitch model ships a TensorFlow.js build, so the whole thing can run in a browser tab. Here is what it took to make that feel responsive, and where it still falls short.

Keep the model out of the first page load

TensorFlow.js plus the model is too heavy for the initial bundle, and most visitors read the page before they drop a file. So the library is only loaded through a dynamic import():

let modulePromise: Promise<typeof import('@spotify/basic-pitch')> | null = null;

function loadModule() {
  if (!modulePromise) {
    modulePromise = import('@spotify/basic-pitch');
  }
  return modulePromise;
}
Enter fullscreen mode Exit fullscreen mode

The model weights (about 1 MB) are served as static files and cached after the first run. The download starts as soon as someone picks a file, so it overlaps with them choosing which part of the audio to convert.

One trap: basic-pitch pins its own TensorFlow.js version. Adding @tensorflow/tfjs as a separate dependency loads a second copy that re-registers every kernel, so don't.

Split inference from note extraction

The model produces three per-frame matrices: note activations (frames), onsets, and pitch contours. Turning those into notes is a separate, cheap step controlled by two thresholds and a minimum note length.

That split matters for the interface. Running the model over a 30 second clip takes a few seconds; turning its output into notes takes milliseconds. So the raw output is kept, and moving a sensitivity slider only re-runs the second step:

// Runs once per segment
const output = await analyze(audioBuffer);

// Runs on every slider change
const notes = await notesFromOutput(output, {
  onsetThreshold,
  frameThreshold,
  minNoteLengthMs,
  minPitchMidi,
  maxPitchMidi,
});
Enter fullscreen mode Exit fullscreen mode

One detail: outputToNotesPoly mutates the frame and onset arrays when you pass a pitch range, so you need to copy them before each call if you want to reuse the originals.

The model expects mono audio at 22050 Hz. OfflineAudioContext handles the mixdown and resampling without any extra library.

Let people hear the mistakes

A piano roll doesn't tell you much about whether the transcription is right. What helped more was an A/B player: the original audio and a simple WebAudio synth of the detected notes play on one shared playhead, and you switch between them. A wrong octave or a missing note is obvious by ear in a second.

Clean up with plain functions

Model output needs tidying before it is useful in a DAW. Each cleanup step is a pure function over a note array, which keeps them easy to test and combine:

  • remove notes shorter than a threshold
  • remove notes much quieter than the rest
  • merge repeated notes separated by tiny gaps
  • trim overlaps so a monophonic line stays monophonic
  • quantize to a grid based on an estimated tempo

Export uses @tonejs/midi to write a standard .mid.

Where it doesn't work

basic-pitch does well on a single line: a sung melody, a guitar riff, a piano part. On a dense full mix it gives you something rough at best. Limiting a run to 15 to 60 seconds also keeps memory in check on phones, where long clips can run out of resources.

The result is MidiDraft, if you want to try it on your own recordings. The examples page includes one clip it deliberately gets wrong, so you can see what failure looks like before trusting it with your own audio.

Top comments (0)