Datasets

Datasets, corpora and benchmarks for music and audio AI.

98 listings

A turquoise atlas field connects tools and organizations around one orange coordinate.

98 results

MagnaTagATune
Magnatune audio clips with human tags collected through the TagATune game for music auto-tagging research.
Global
MARBLE
A unified evaluation suite for pretrained music-audio representations across acoustic, performance, score and high-level tasks.
Global
MedleyDB
Annotated royalty-free multitrack recordings supporting melody, pitch, instrument and source-separation research.
United States
MID-FiLD
Pozalabs MIDI dataset of 4,422 professional-writer samples for fine-level dynamics control.
South Korea
Million Song Dataset
Audio features and metadata for one million contemporary popular-music tracks, without distributed audio.
United States
MIR-1K
Dataset of 1,000 Chinese karaoke song clips with separated accompaniment/vocal channels and manual annotations.
Taiwan
MMAU
A multi-domain benchmark of expert audio understanding and reasoning spanning speech, environmental sound and music.
United States
MMAU-Pro
Expert-created audio questions covering speech, sound, music, mixtures, spatial audio and long-form reasoning.
United States
MoisesDB
Multitrack music dataset with a hierarchical stem taxonomy for fine-grained source-separation research.
Brazil
MuChoMusic
Human-validated multiple-choice questions for evaluating music understanding and reasoning in audio-language models.
Global
MUSDB18
Full-length stereo music mixtures with isolated drums, bass, vocals and other stems for source separation.
Global
MUSIB
Open benchmark and software package for reproducible evaluation of musical-score inpainting systems.
Chile
Music Arena
A live pairwise evaluation platform and rolling leaderboard for text-to-music generation systems.
United States
Music Arena Dataset
Rolling human preference data from live pairwise comparisons of text-to-music model outputs.
United States