Disclosure: I work on NeuralSound, one of the three tools tested here. The method, the raw numbers and the caveats are all below, so you can judge them yourself or rerun the test.
Every AI vocal remover says it's "studio quality". Almost none of them publish numbers. So we ran a small, reproducible benchmark against two popular tools, Moises and Fadr, and scored all three with standard source-separation metrics.
This post covers:
- What SI-SDR, SI-SIR and SI-SAR measure, and why they're more useful than "sounds good to me"
- The exact test setup, plus Python you can reuse
- The results, and what they do and don't prove
- Where separation fits in a real workflow: background music removal, karaoke, and lyrics in 96+ languages
The problem: "sounds clean" isn't a metric
Music source separation takes one mixed stereo track and splits it into stems: vocals, drums, bass, everything else. Whether you want a vocal remover (to keep the instrumental), a vocal isolator (to keep the a cappella) or a background music remover (to keep speech in a video), it's the same underlying task.
Listening tests matter, but they're subjective and hard to compare. Objective metrics give you a repeatable baseline.
The metrics, in plain English
All three are measured in decibels, and higher is better.
| Metric | Question it answers |
|---|---|
| SI-SDR (Scale-Invariant Signal-to-Distortion Ratio) | Overall, how close is the separated stem to the true stem? |
| SI-SIR (Signal-to-Interference Ratio) | How much of the other source leaked in? (e.g. hi-hats left in the vocals) |
| SI-SAR (Signal-to-Artifact Ratio) | How many new artifacts did the model add? (that "watery" or robotic sound) |
"Scale-invariant" means a stem that's simply quieter than the reference isn't penalized. Only its shape matters.
Here's the core of SI-SDR in NumPy:
import numpy as np
def si_sdr(est: np.ndarray, ref: np.ndarray, eps: float = 1e-8) -> float:
"""Scale-invariant SDR in dB. est and ref are 1-D mono signals of equal length."""
est = est - est.mean()
ref = ref - ref.mean()
alpha = np.dot(est, ref) / (np.dot(ref, ref) + eps) # optimal scaling
target = alpha * ref # part of est explained by ref
noise = est - target # everything else
return 10 * np.log10((np.sum(target**2) + eps) / (np.sum(noise**2) + eps))
Every 3 dB is roughly a 2× reduction in error energy, so small-looking gaps are bigger than they seem.
Test setup
We kept it simple and easy to reproduce:
- Dataset: the first five songs in the valid split of MUSDB18-HQ (uncompressed, with isolated ground-truth stems)
- Task: two-stem separation (vocals + instrumental)
- Processing: every output was converted to mono at 44.1 kHz and aligned to the reference with cross-correlation, because different services add different padding or latency
- No preprocessing of any kind before scoring
The alignment step is easy to skip, and skipping it quietly wrecks your scores:
from scipy.signal import correlate
def align(est: np.ndarray, ref: np.ndarray):
"""Shift est so it lines up with ref, then trim both to the same length."""
corr = correlate(est, ref, mode="full", method="fft")
lag = int(np.argmax(corr)) - (len(ref) - 1)
if lag > 0:
est = est[lag:]
elif lag < 0:
est = np.pad(est, (-lag, 0))
n = min(len(est), len(ref))
return est[:n], ref[:n]
Results
Averages across the five songs:
| Tool | Overall SI-SDR | Vocals SI-SDR | Instrumental SI-SDR | SI-SIR | SI-SAR |
|---|---|---|---|---|---|
| NeuralSound | 15.80 dB | 13.26 dB | 18.33 dB | 33.30 dB | 15.93 dB |
| Moises | 14.47 dB | 12.00 dB | 16.94 dB | 30.25 dB | 14.64 dB |
| Fadr | 12.59 dB | 10.02 dB | 15.16 dB | 25.55 dB | 12.89 dB |
Full breakdown and methodology: AI Vocal Remover Benchmark 2026.
What the numbers say
- Instrumentals score higher than vocals for every tool. That's expected. The instrumental is most of the mix's energy, so small vocal leftovers barely dent its score. Vocals are the harder half of the problem.
- The biggest gap is in SI-SIR (interference). NeuralSound's 33.3 dB against Fadr's 25.6 dB is about 7.7 dB, which is roughly 6× less leakage energy. In practice, that's the difference between "faint drums behind the singer" and "clean a cappella".
- SI-SAR is closer across the board. All three models add some artifacts. Nobody has "solved" separation.
What the numbers don't say
- Five songs is a small sample. It's a useful signal, not a universal ranking.
- MUSDB18 is mostly Western pop and rock. Heavily reverberant, live or lo-fi recordings behave differently.
- Metrics aren't taste. A model can score well and still sound slightly thin on a particular track. Always listen too.
If you rerun this on other songs, I'd love to see your numbers in the comments.
From benchmark to real use cases
The same separation engine powers several everyday tasks:
- Background Music Remover: remove background music from vlogs, interviews, lectures and podcasts. Speech stays, music goes. The quickest way to remove music from audio.
- Vocal Remover: strip vocals to make instrumentals and karaoke tracks.
- Instrumental Remover: remove the instrumental from a song and keep just the singer.
- Vocal Isolator: get the cleanest possible a cappella for remixes or transcription.
- Music Separation: split a track into vocals, drums, bass and other stems.
Bonus: lyrics in 96+ languages, under 5% error rate
Once vocals are isolated, transcription gets much easier, because the speech model isn't fighting drums and synths. We use that for automatic lyrics and captions:
- 96+ languages supported
- Under 5% lyrics text error rate in our internal tests
That feeds the Karaoke Video Maker (synced on-screen lyrics) and Auto Caption Generation (subtitles for videos after you remove the music).
Try it
- Web: neuralsound.org
- Android: NeuralSound on Google Play
- iOS: NeuralSound AI Vocal Remover on the App Store
Questions, bug reports or benchmark suggestions: drop a comment or email support@neuralsound.org.
What would you want added to the next round: more songs, more tools (Demucs, UVR, LALAL.AI), or a blind listening test?

Top comments (0)