Ground Truth.
AI, checked against the source.

News · 2026-07-27

Leaked Suno code names YouTube, Deezer, Genius and podcast RSS feeds as collection sources

A hack of the AI music generator Suno has exposed source-code files and comments naming YouTube Music, Deezer, Genius, stock music libraries and podcasts via RSS feeds as data collection sources. 404 Media, which examined the material, reports it as the most specific public evidence yet about where a major music model's training data came from - arriving while Suno is in litigation with the recording industry over exactly that question.

Key facts

Why provenance is the whole fight

Suno has never claimed its models learned from licensed catalogues. In its court answer it stated the model learned from tens of millions of recordings from publicly available sources, and its chief executive has separately said it trains on medium- and high-quality music found on the open internet. Its public position frames this as statistical learning rather than copying.

The RIAA's landmark cases call it unlicensed copying at scale. Both sides have argued about a corpus neither side has publicly enumerated.

What changes with this leak is specificity. "Publicly available sources" is a legal abstraction. A code comment naming Deezer is a collection decision with a target, and pipeline-level labels are the kind of artifact that turns an inference into a factual dispute with documents attached. That is a meaningful shift in a case that has run largely on expert argument about how models memorize.

The podcast angle, stated precisely

Podcast RSS feeds appear as a named category in the reported material. That is genuinely notable - podcast audio is speech, not music, and its inclusion suggests a collection pipeline casting wider than the product's stated purpose. RSS is also the easiest bulk audio source on the internet: an open, unauthenticated feed of direct download links, published by design.

But the evidence stops at the category. Nothing public names a show, an episode or a feed, and nothing establishes that podcast audio reached a released Suno model. A collection pipeline is not a training set, and a training set is not a shipped model. Anyone reporting that their podcast was used to train Suno is going beyond what the material supports.

The security dimension

The mechanism here deserves attention independent of the copyright fight, because it is becoming a pattern. The most consequential thing exposed in an AI company breach is increasingly not customer records - it is the pipeline: what was collected, from where, and with what code. Model weights, training corpora, scraping infrastructure and evaluation harnesses are now the crown jewels, and they are documented in ordinary repositories with ordinary access controls.

That is the same shift visible in the JadePuffer campaign, where a follow-on tool was purpose-built to destroy AI artifacts - model checkpoints, vector indexes, training data. Attackers and leakers have both worked out that the valuable thing in an AI company is the data supply chain. Our lesson on training data deduplication covers why what goes into that pipeline shapes model behavior so directly, and synthetic data covers the main alternative labs reach for when scraping gets legally expensive.

The honest caveat

Everything above rests on one outlet's examination of leaked material. There is no published corpus, no hash-verified code archive, no model-to-file audit, and no independent second examination of the raw code found in this pass. The reported sources may also be partial - nothing establishes they represent Suno's whole training corpus, and Suno characterizes the underlying incident as involving obsolete code, which if accurate would mean the pipeline described is not necessarily the current one.

Suno's full statement, including its account of the November 2025 incident, is carried in Pitchfork's report.

The right way to hold this story is the way 404 Media reported it. This is not a leaked playlist of the songs and episodes inside Suno's models. It is reportedly a collection pipeline, naming podcasts via RSS alongside music sources, in code the company says is out of date. That is important provenance evidence and a genuinely new kind of document in the AI copyright fight. It is not an auditable account of what any current model learned from, and the difference between those two things is exactly where the litigation will be fought.


Primary source, verified: read the paper →

Key questions

What did the Suno leak actually reveal?

404 Media, which examined the material, reports that leaked source-code files and comments name YouTube Music, Deezer, Genius, stock music libraries and podcasts via RSS feeds as data collection sources. This is pipeline-level evidence, not a published corpus or a model-to-file audit.

Does this prove podcasts were used to train Suno's models?

No. Podcast RSS feeds appear as a named category in the collection code, but there is no public evidence naming specific shows or episodes, and nothing establishing that any podcast audio reached a released Suno model.

What has Suno said about its training data?

Suno had already stated in its court answer that its model learned from tens of millions of recordings from publicly available sources, and its chief executive has said it trains on medium- and high-quality music found on the open internet. On the breach itself, Suno says it contained a limited November 2025 incident involving obsolete code and that no sensitive personal data was compromised.
Cite this

APA

Ground Truth. (2026, July 27). Leaked Suno code names YouTube, Deezer, Genius and podcast RSS feeds as collection sources. Ground Truth. https://groundtruth.day/news/suno-leak-names-podcast-rss-feeds-among-training-sources.html

BibTeX

@misc{groundtruth:suno-leak-names-podcast-rss-feeds-among-training-sources,
  title  = {Leaked Suno code names YouTube, Deezer, Genius and podcast RSS feeds as collection sources},
  author = {{Ground Truth}},
  year   = {2026},
  month  = {jul},
  url    = {https://groundtruth.day/news/suno-leak-names-podcast-rss-feeds-among-training-sources.html}
}

Topics: cybersecurity · data-breach · copyright · training-data · music · ai-provenance

Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.