Datasets

Datasets, corpora and benchmarks for music and audio AI.

98 listings

A turquoise atlas field connects tools and organizations around one orange coordinate.

98 results

CMI-Bench
A test-only benchmark converting diverse MIR tasks into standardized music instruction-following evaluation.
Global
CMI-RewardBench
Preference data and evaluation for reward models scoring musicality, text alignment and compositional instruction following.
Global
CocoChorales
Synthetic four-part chamber-ensemble music with mixtures, stems, MIDI and fine-grained performance annotations.
United States
ComMU
Pozalabs symbolic dataset of 11,144 professional-composer MIDI samples with 12 metadata fields.
South Korea
CoversBR
Large predominantly Brazilian music database for cover-song and live-song identification, with public metadata and features.
Brazil
CtrSVDD
Controlled singing-voice deepfake corpus generated from public singing data with synthesis and voice-conversion systems.
Global
DALI
Full-song lyric, vocal-note and audio alignments at note, word, line and paragraph granularity.
France
DEAM
Music excerpts and full songs with continuous and song-level valence and arousal annotations.
Europe
EgoMusic
Synchronized egocentric audiovisual dataset for music understanding and hearing-enhancement research.
Ireland
EMOPIA
Pop-piano audio and MIDI clips labeled for perceived emotion, supporting recognition and emotion-conditioned generation.
Taiwan
Erkomaishvili Dataset
Curated and annotated corpus of traditional Georgian vocal music for computational musicology and MIR tasks.
Georgia / Georgian music
FakeMusicCaps
MusicCaps prompts regenerated with multiple text-to-music systems for synthetic-music detection and generator attribution.
Italy
Free Music Archive Dataset
Full-length Creative Commons music with audio, features, metadata and hierarchical genre labels for MIR.
Global