The Atlantic launched a searchable database for the music employed in AI training.

The Atlantic launched a searchable database for the music employed in AI training.
Summary
Alex Reisner uncovered four music datasets for AI training, with two having millions of tracks.
Sets have been downloaded thousands of times; Google and Stability confirmed their use in research.
Datasets are complex to use; some methods violate terms of service for platforms like YouTube.

Share

Bookmark

Newsletter

The Verge's weekend editor, with over 18 years of experience in the tech industry and a deep understanding of synthesizers, has a new perspective on the intersection of music and AI.

In a recent discovery by Atlantic's Alex Reisner, four significant music datasets designed for training AI models are now accessible to the public. Among these, two massive collections boast an impressive 12 million and 9 million tracks, while the remaining two, though smaller, still contribute valuable data with over 100,000 songs each.

Reisner notes that these datasets have been downloaded thousands of times. Although it’s challenging to track who utilizes them, both Google and Stability have acknowledged their use in various research publications. Some of these datasets, like those from the Free Music Archive, can be streamed for personal use without charge but require appropriate licensing for any commercial endeavors.

Despite their online availability, utilizing these datasets for AI training isn't as straightforward as simply downloading a ZIP file. Reisner elaborates on the process: three of the datasets consist of lists linking to songs available on platforms like YouTube or Spotify. AI developers typically use automated tools to download the audio files, which often circumvent necessary logins and advertisements. This method raises concerns, as it violates the terms of service of the platforms involved.

Loading comments...