DEV Community

Priyansh Kansara
Priyansh Kansara

Posted on

I rejected a model that passed everyone else's benchmark

I vendored an audio deepfake detector into my verification platform. It loaded cleanly, every integration check passed, and the forward pass looked healthy. Then I measured it on a held-out slice and got an AUC of 0.54 — barely above chance. The vendor reports 0.9967. This is the story of what happened in between, and why I decided not to ship it.

The component that worked

The model is Spectra-AASIST3, an audio deepfake detector released under Apache-2.0. I vendored it into the project as a potential audio channel for a WhatsApp-first deepfake and malware verification platform. The integration work was straightforward and, by every check I ran, successful.

The checkpoint loaded without a single mismatch: 1022 out of 1022 keys matched and applied, 318.9 million parameters. The forward pass was responsive and produced sensible outputs across a range of inputs. I probed it with silence, real speech, envelope noise, and a pure tone. The bonafide logit came back at 1.63 on silence, 1.76 on speech, 1.33 on envelope noise, and 2.8 to 4.0 on tone. The model was clearly doing something — it was not broken, not degenerate, not returning constant outputs. It was alive and responsive.

By every integration test I had, this component worked.

But an integration test asks a narrow question: does the component run? Does it load? Does it produce outputs? Does it respond to inputs? It does not ask the question that actually matters: does it discriminate? Can it tell a real voice from a synthetic one on the data I care about? Those are different questions, and conflating them is how you end up shipping a model that passes every unit test and fails every user.

The bar, fixed before measuring

Before I ran the evaluation, I fixed the acceptance criterion in writing: AUC ≥ 0.85 on a held-out slice, with non-overlapping posterior quantiles between classes and a global posterior standard deviation above 0.01. Three conditions, all of which had to hold.

Fixing the bar in advance is the only reason this story has a clean ending. Once you have a number in front of you — any number — the temptation is to reason backwards. A 0.54 can become "not bad for a small slice" or "the bar was too strict" or "the vendor's number is on a different set anyway." All of those things might be true, and none of them matter if the bar was committed before the measurement. The bar is not a conclusion. It is a decision rule. You set it when you are still honest, and you live with the result.

The measurement

I evaluated on a 240-clip slice of ASVspoof 2019 LA — 120 real, 120 fake — drawn from the Bisher validation parquet. The result: AUC 0.54. The posteriors told the same story more starkly. Both class 0 and class 1 had interquartile ranges sitting at roughly 0.044 to 0.049. The distributions did not just overlap; they were essentially on top of each other. The global posterior standard deviation was 0.0046, an order of magnitude below the 0.01 threshold I had set.

The model was producing confident outputs. They were just not different outputs for real versus fake audio.

The gap I cannot explain

The vendor's published AUC on ASVspoof 2019 LA is 0.9967. My measured number is 0.54. That is not a small gap, and I want to be honest about what I know and do not know.

The vendor's number is computed on the full test set — 71,237 trials. My measurement is on a 240-clip slice of the validation parquet. These are not the same evaluation set, and the difference in scale is enormous. It is possible that the model performs dramatically better on the full test distribution than on this particular slice. It is also possible that something about the slice — its provenance, its label balance, its acoustic characteristics — makes it harder than the test set. I do not know which of these is true, and I am not going to pretend I do.

What I can say is that the discrepancy is unresolved. Resolving it would require re-running the evaluation on the official ASVspoof 2019 LA evaluation audio, which is 2.7 GB of data I have not yet processed. Until that re-proof happens, the 0.54 stands as the only measurement I have on data I can defend, and it is the number I have to make the decision on.

The decision

I did not wire the model into the serving path. Audio stays calibration-only. The vendored code is committed in-tree so that it can be re-proven quickly if official-eval evidence arrives, but the base-encoder weights are kept out of git, and the model is not used in any user-facing decision.

The project does not need an audio channel that does not work. It already has a fusion engine that combines calibrated per-signal evidence across video, image, URL, and government document channels. Adding a component that scores 0.54 would not make the system better. It would make it more complex and less trustworthy, because every downstream consumer would have to account for a signal that carries no information.

The leakage finding

This was not the first time I had to make this kind of call. Earlier in the same project, I trained a logistic regression fusion stack on out-of-fold predictions from the base audio detectors. The out-of-fold AUC was 0.512, against a heuristic baseline of 0.502. A gradient-boosted alternative scored 0.499 at n=84 — chance. Neither was wired.

The uncomfortable part of that story is what the single-split training run reported: 0.634. That number was real, in the sense that it was a genuine output of a genuine training run. It was also flattering, because the split happened to favor the model. The out-of-fold number — 0.512 — was the honest one, and it was the worse one. I reported the worse number.

This is a discipline that is easy to state and hard to practice. When you have two numbers and one of them is better, the better one is the one that wants to be in the commit message, the README, the portfolio. The worse one is the one that reflects what the model will actually do on data it has not seen. You report the worse number, and you explain why it is the worse number, and you move on.

The lesson

An integration test proves that a component runs. It does not prove that the component discriminates. Loading a checkpoint, matching keys, checking that the forward pass responds to inputs — these are necessary conditions, and they are easy to satisfy. They are also the conditions that every vendor demo and every quick-start notebook satisfies, because they are the conditions that are easy to show.

Discrimination is a separate property, and it has to be measured on data that represents the problem you actually have, against a bar you set before you see the result. If you do not fix the bar in advance, you will rationalize whatever number appears. You will find reasons why the slice is hard, why the vendor's number is on a different set, why 0.54 is close enough. You will be wrong, and you will not know it, because you will have convinced yourself after the fact.

I would rather ship a system with fewer channels and more confidence in the ones it has than a system with an audio detector that passes every integration test and cannot tell a real voice from a fake one. The model was not broken. It was rejected by a bar I set before I measured it. That is the whole story, and it is the interesting part.

Top comments (0)