You can turn a short hummed recording into MIDI in the browser with Spotify's Basic Pitch. Decode the audio, resample it to mono at 22,050 Hz, run the model, then encode its detected notes. The result needs a listen and usually some editing.
I build SoundToMIDI around that workflow. One design decision turned out to matter as much as the model: moving inference into a Web Worker so the interface could respond while a conversion was running.
The use case is familiar. In this r/hardstyle thread, a producer describes recording a melody, converting it, then spending too long correcting the notes. That is the job here: get an idea into a DAW without entering every note by hand.
From recording to note events
Local file
→ Web Audio decoding
→ mono PCM at 22,050 Hz
→ Basic Pitch inference
→ onset/frame post-processing
→ note events in seconds
→ MIDI encoding
→ piano-roll inspection and download
Basic Pitch supports polyphonic notes, though Spotify recommends one instrument at a time. Start with a solo hum. MIDI records note events; it won't preserve the sound of your voice.
Let AudioContext decode at the device's supported rate, then use a mono OfflineAudioContext to prepare the model input. Forcing the decoding context to 22,050 Hz can cause problems on some devices.
async function decodeAndResample(file) {
const context = new AudioContext();
let decoded;
try {
decoded = await context.decodeAudioData(await file.arrayBuffer());
} finally {
await context.close();
}
const sampleRate = 22050;
const offline = new OfflineAudioContext(
1,
Math.max(1, Math.ceil(decoded.duration * sampleRate)),
sampleRate,
);
const source = offline.createBufferSource();
source.buffer = decoded;
source.connect(offline.destination);
source.start();
const rendered = await offline.startRendering();
return Float32Array.from(rendered.getChannelData(0));
}
This applies Web Audio's channel-mixing rules. Stereo phase cancellation can affect the mono result. A mono recording avoids that surprise. File extensions can be misleading too: a supported container may hold a codec the browser cannot decode.
For a prototype in a browser project with a bundler such as Vite, install:
npm install @spotify/basic-pitch@1.0.1 @tonejs/midi@2
Copy the model assets from basic-pitch-ts into public/model/, preserving the paths between model.json and its weight shards. Add this to the same module:
import {
BasicPitch,
outputToNotesPoly,
noteFramesToTime,
} from '@spotify/basic-pitch';
import { Midi } from '@tonejs/midi';
const engine = new BasicPitch('/model/model.json');
export async function fileToMidi(file, onProgress = () => {}) {
const pcm = await decodeAndResample(file);
const frames = [];
const onsets = [];
await engine.evaluateModel(
pcm,
(f, o) => {
frames.push(...f);
onsets.push(...o);
},
onProgress,
);
const minFrames = Math.max(1, Math.round(0.130 * (22050 / 256)));
const events = noteFramesToTime(
outputToNotesPoly(frames, onsets, 0.5, 0.3, minFrames),
);
if (!events.length) throw new Error('No pitched notes detected.');
const midi = new Midi();
midi.header.setTempo(120);
const track = midi.addTrack();
track.name = 'Transcribed melody';
for (const note of events) {
track.addNote({
midi: note.pitchMidi,
time: note.startTimeSeconds,
duration: Math.max(0.02, note.durationSeconds),
velocity: Math.max(0, Math.min(1, note.amplitude)),
});
}
return { bytes: midi.toArray(), events };
}
Pass a File from an <input type="file"> to fileToMidi. Turn its bytes into an audio/midi Blob for download, and revoke the object URL when finished.
The 0.5 onset threshold controls note starts; 0.3 controls sustained pitch activity. The minimum duration is in frames, so 130 ms becomes roughly 11 at 22050 / 256 frames per second. Raising these controls may remove breath-triggered blips, but can also lose quiet or short intended notes. Change one setting and compare the same phrase.
This prototype runs on the main thread. It is enough to inspect a short clip, but the application needs a different execution boundary.
Why an async conversion can still freeze Cancel
Awaiting a Promise doesn't move synchronous inference kernels off the main thread. SoundToMIDI keeps decoding and native resampling outside the Worker, and puts TensorFlow, inference, note extraction, and MIDI encoding inside it.
The boundary passes PCM and settings:
const worker = new Worker(
new URL('./transcription-worker.js', import.meta.url),
{ type: 'module' },
);
// After a separate model-load request has completed:
worker.postMessage(
{
id: 42,
type: 'transcribe',
mono: pcm,
options: { onsetThreshold: 0.5, frameThreshold: 0.3 },
},
[pcm.buffer],
);
Transferring pcm.buffer avoids a copy and detaches it from the sender. Keep the decoded audio if the user may retry.
Cancel terminates the Worker and rejects pending requests. Retry creates a new one, which must load the model again. A cancellation flag alone can't interrupt a kernel that hasn't returned.
There is a race to test here: cancel, retry immediately, then receive a late response from the old Worker. Check the Worker instance as well as the request ID before accepting a response. An obsolete result must not complete the new request.
The memory bill follows the recording
Basic Pitch 1.0.1 uses overlapping windows. The application's inference loop processes one window at a time and releases its temporary tensors after reading the values. Leading padding and overlap trimming must match the reference implementation, or window boundaries can alter the output.
Use tf.tidy for synchronous tensor work, then dispose returned tensors in finally after awaiting their arrays. An async callback inside tf.tidy won't manage tensor lifetime across an await.
The decoded audio and accumulated frame outputs still grow with clip length. Windowing doesn't make this a constant-memory streaming system. I prefer short selections for this workflow.
The model runs locally, but the browser downloads code and weights. Local audio processing doesn't make the page network-free.
Eight intended notes can become nine
On October 8, 2026, we reran the published synthetic fixtures in desktop Chrome using the live converter. These are eight-second WAVs at 22,050 Hz. Settings were onset 0.5, frame 0.3, minimum length 130 ms, tempo 120 BPM, and pitch bends off:
| Input | Intended pitched events | Raw output events | Extra events |
|---|---|---|---|
| Isolated melody | 8 | 9 | 1 |
| Chords with noise percussion | 12 | 15 | 3 |
| Unpitched noise percussion | 0 | 0 | 0 |
Every pitched source event had a same-pitch onset match within 150 ms, used once. We didn't score duration or expression. These synthetic fixtures cannot establish accuracy on humming. The noise-only clip produced no downloadable MIDI.
The melody's extra D5 was an octave above a source D4. Harmonic confusion is a possible explanation; the octave relationship doesn't prove the cause. The source audio and original September 28 results are public. That earlier melody run produced ten events, which is another reason to record the test date and settings.
Check repeated notes and phrase endings when you try your own recording. Also remember that writing 120 BPM into a MIDI header doesn't detect the recording's beat grid. This code writes notes in seconds against a chosen tempo map.
I leave pitch bends out of the prototype. Ordinary MIDI bends affect a channel, so overlapping notes can't all bend independently on a single channel. Cleanup rules can shorten overlaps or remove blips, but they don't know the intended melody. Keep the raw events available.
Record 10–20 seconds without backing music, leave gaps between repeated notes, and try the humming-to-MIDI workflow. Raw MIDI downloads are free; optional paid exports are separate. Listen beside the piano roll before importing it into a DAW. That last check catches errors that a successful download won't.
Top comments (2)
The onset-only score leaves a useful blind spot for the overlap-trimming test: a note could start within 150 ms yet end too early at a window boundary, or sustain across an intended gap. Have you tried the same synthetic phrase shifted so a repeated-note gap falls just before, on, and just after that boundary?
I'd compare raw events against the reference for onset, offset and whether the repeated note stays split, including a final note that reaches the clip end. That would exercise the window assembly separately from the humming-accuracy question. I haven't run SoundToMIDI; this is a suggested continuation of the synthetic fixtures, not a claim of a trimming bug.
Thanks, that’s a useful distinction. The 150 ms check only scores pitch/onset matches, so it doesn’t tell us whether note endings or gaps survived.
The current code joins the trimmed frame/onset outputs before extracting notes. I also have assembly tests against Basic Pitch 1.0.1, but those use simulated model outputs rather than an actual repeated-note phrase. I haven’t run the shifted-gap cases yet.
Comparing onset, offset, split/merged repeats, and the clip-end note against both the known event schedule and the reference implementation would be a good next test. Thanks for spelling out the cases.