Leaked source code from Suno, the AI-powered music generation startup, has revealed the extent of the data used to train its models. According to the leak, the company fed its artificial intelligence system with thousands of hours of copyrighted music from platforms including Deezer, YouTube, and Pond5.
The exposed source code provides detailed insight into how Suno assembled its training library. The documents show that the startup scraped a vast collection of audio files from these streaming services, using them to develop its generative music capabilities that allow users to create songs from text prompts.
The leak comes amid ongoing legal battles between AI music companies and record labels over the use of copyrighted material for training purposes. Suno, along with competitor Udio, has faced lawsuits from major recording industry organizations alleging mass copyright infringement. The companies have argued that their use of publicly available music constitutes fair use under copyright law, similar to how humans learn from existing works.
The leaked data suggests that Suno’s training approach was more extensive than previously understood, drawing from commercial streaming platforms rather than just publicly available datasets. Critics argue that this constitutes “staggering theft” of artists’ work, while supporters of the technology maintain that the training methods are consistent with how other AI systems have been developed.
Suno has not officially confirmed the authenticity of the leaked code or commented on the specific training sources detailed in the documents. The company previously acknowledged in court filings that it had collected music from the open internet but did not specify which platforms were used.
This article was adapted from Decrypt. Read the original here.
