A hack of the AI music generator Suno has exposed source-code files and comments naming YouTube Music, Deezer, Genius, stock music libraries and podcasts via RSS feeds as data collection sources. 404 Media, which examined the material, reports it as the most specific public evidence yet about where a major music model's training data came from - arriving while Suno is in litigation with the recording industry over exactly that question.
Key facts
- Named collection sources in the leaked material include YouTube Music, Deezer, Genius, stock libraries and podcast RSS feeds.
- Source-code files and comments, not output-based inference - 404 Media examined the material directly.
- Suno says the underlying incident dates to November 2025, was limited and contained, involved obsolete code, and compromised no sensitive personal data.
- Primary source: 404 Media's report.
Why provenance is the whole fight
Suno has never claimed its models learned from licensed catalogues. In its court answer it stated the model learned from tens of millions of recordings from publicly available sources, and its chief executive has separately said it trains on medium- and high-quality music found on the open internet. Its public position frames this as statistical learning rather than copying.
The RIAA's landmark cases call it unlicensed copying at scale. Both sides have argued about a corpus neither side has publicly enumerated.
What changes with this leak is specificity. "Publicly available sources" is a legal abstraction. A code comment naming Deezer is a collection decision with a target, and pipeline-level labels are the kind of artifact that turns an inference into a factual dispute with documents attached. That is a meaningful shift in a case that has run largely on expert argument about how models memorize.
The podcast angle, stated precisely
Podcast RSS feeds appear as a named category in the reported material. That is genuinely notable - podcast audio is speech, not music, and its inclusion suggests a collection pipeline casting wider than the product's stated purpose. RSS is also the easiest bulk audio source on the internet: an open, unauthenticated feed of direct download links, published by design.
But the evidence stops at the category. Nothing public names a show, an episode or a feed, and nothing establishes that podcast audio reached a released Suno model. A collection pipeline is not a training set, and a training set is not a shipped model. Anyone reporting that their podcast was used to train Suno is going beyond what the material supports.
The security dimension
The mechanism here deserves attention independent of the copyright fight, because it is becoming a pattern. The most consequential thing exposed in an AI company breach is increasingly not customer records - it is the pipeline: what was collected, from where, and with what code. Model weights, training corpora, scraping infrastructure and evaluation harnesses are now the crown jewels, and they are documented in ordinary repositories with ordinary access controls.
That is the same shift visible in the JadePuffer campaign, where a follow-on tool was purpose-built to destroy AI artifacts - model checkpoints, vector indexes, training data. Attackers and leakers have both worked out that the valuable thing in an AI company is the data supply chain. Our lesson on training data deduplication covers why what goes into that pipeline shapes model behavior so directly, and synthetic data covers the main alternative labs reach for when scraping gets legally expensive.
The honest caveat
Everything above rests on one outlet's examination of leaked material. There is no published corpus, no hash-verified code archive, no model-to-file audit, and no independent second examination of the raw code found in this pass. The reported sources may also be partial - nothing establishes they represent Suno's whole training corpus, and Suno characterizes the underlying incident as involving obsolete code, which if accurate would mean the pipeline described is not necessarily the current one.
Suno's full statement, including its account of the November 2025 incident, is carried in Pitchfork's report.
The right way to hold this story is the way 404 Media reported it. This is not a leaked playlist of the songs and episodes inside Suno's models. It is reportedly a collection pipeline, naming podcasts via RSS alongside music sources, in code the company says is out of date. That is important provenance evidence and a genuinely new kind of document in the AI copyright fight. It is not an auditable account of what any current model learned from, and the difference between those two things is exactly where the litigation will be fought.
Originally published on Ground Truth, where every claim is checked against the primary source.
Top comments (0)