Leaked Suno Source Code Reportedly Shows Large-Scale Music Scraping from YouTube Music, Deezer, Genius and Pond5

The debate over AI training data has entered a new phase after leaked internal source code from Suno reportedly revealed detailed scraping operations targeting major music and lyrics platforms, including YouTube Music, Deezer, Genius, and commercial stock-music marketplace Pond5.

The leaked materials, obtained through a security breach and later reviewed by journalists, appear to contain scraping scripts, dataset references, and ingestion volumes that allegedly document how audio was collected for Suno’s AI training models. If authentic, the documents provide the most detailed evidence to date of the specific sources used to build the company’s music-generation system.

Why the leak matters

For more than a year, major record labels and industry groups have argued that AI music companies trained their systems on copyrighted recordings without obtaining licenses. Those accusations formed a central part of the RIAA lawsuit filed against Suno in 2024.

Until now, many of those claims were based on legal allegations, technical analysis, and public statements from the company. The leaked code reportedly changes the conversation by identifying particular platforms and showing the scale of the data collection.

According to the report, Suno ingested more than 100,000 hours of audio from YouTube Music alone, alongside substantial amounts of material from other services. In previous court filings, Suno acknowledged that it trained on “essentially all music files of reasonable quality accessible on the open internet,” but the leak reportedly provides a clearer picture of what that meant in practice.

The Pond5 revelation could affect independent producers

One of the most significant details for electronic producers and independent artists is the reported inclusion of Pond5.

Pond5 operates as a commercial licensing marketplace where musicians upload tracks specifically to be licensed for film, advertising, games, and other media projects. If those recordings were scraped for AI training without additional permission, many independent creators could find themselves in a particularly unusual position: music they licensed through a legitimate marketplace may have been used to train systems capable of generating competing music.

The concern arrives amid growing scrutiny of AI training datasets. A recent investigation by The Atlantic reported that more than 21 million tracks from well-known artists, including Flume, Tame Impala, and Sia, had been used in AI training datasets without authorization.

Suno’s response

Suno confirmed that a security incident occurred in November 2025 but said it was quickly contained and involved “outdated source code no longer in use.” The company stated that it did not believe individual customer notifications were required under applicable privacy laws.

On the copyright issue, Suno continues to argue in ongoing litigation that training AI models on copyrighted material qualifies as fair use, a legal defense that remains unresolved in courts.

The company has not publicly confirmed the authenticity of every leaked file cited in the reporting.

A wider crisis for the music industry

The disclosure arrives at a moment when AI-generated music is becoming impossible to ignore. Industry analysts estimate that around 120,000 new tracks are uploaded to streaming services every day, with AI-generated content accounting for a rapidly growing share of that volume.

For producers, DJs, and labels, the issue is no longer theoretical. AI systems trained on existing music are increasingly capable of generating tracks that resemble established genres, production techniques, and stylistic signatures. The concern among many creators is not only about compensation for training data, but also about visibility: human-made releases must now compete against algorithmically generated music in the same recommendation systems and streaming feeds.

The unresolved question

The leaked code may strengthen the evidentiary basis for ongoing lawsuits, but it does not by itself determine the legal outcome. Courts will still need to decide whether large-scale scraping and AI training on copyrighted recordings without explicit permission constitutes fair use, copyright infringement, or another category entirely.

What is already clear is that the leak has intensified pressure on AI music companies and renewed calls for clearer regulation of training data practices. For the electronic music community, the controversy cuts to the heart of a fundamental question: who owns the sounds that future AI systems are built upon?

Leave a Reply

Your email address will not be published. Required fields are marked *