Datasets

Datasets, corpora and benchmarks for music and audio AI.

98 listings

A turquoise atlas field connects tools and organizations around one orange coordinate.

98 results

SAMBASET
Dataset of historical samba-enredo recordings created for MIR research on Brazilian music.
Rio de Janeiro, Brazil
Saraga collections
Open annotated audio corpora for Carnatic and Hindustani music research.
India
SingFake
In-the-wild real and deepfake singing data for singing-voice authenticity detection.
United States
Slakh2100
Synthetic multitrack audio and aligned MIDI rendered from Lakh MIDI files for separation and transcription.
United States
Song Describer Dataset
Crowdsourced human descriptions of Creative Commons music for music-language evaluation.
Global
SongCompose-PT
Chinese-English pretraining dataset containing lyrics, melodies and paired lyric-melody data for SongComposer.
Mainland China, China
SongEval
Full-length real and generated songs with expert ratings across five aesthetic dimensions.
China
Sounds Queer
Dataset of AI-generated music from queer-identity prompts for auditing text-to-music models.
Germany
Suno Music Generation Dataset
Metadata and downloadable links for 659,788 Suno-generated songs discovered through systematic search queries.
Global
TinySOL
Compact dataset of isolated instrumental notes published through IRCAM Forum.
Paris, France
Tohoku Kiritan Singing Database
東北きりたん歌唱データベース
Research singing database used to create the Tohoku Kiritan library distributed with NEUTRINO.
Japan
Tunepal tune corpus
Search corpus of more than 23,000 Irish traditional music tunes used by Tunepal.
Ireland
URMP Dataset
Coordinated multi-instrument classical performances with separated audio, assembled video, scores and pitch annotations.
United States
VocalSet
A cappella singing recordings across singers, vowels, registers and vocal techniques, with a corrected annotated derivative.
United States