FSD50KOpen human-labeled Freesound clips across 200 AudioSet-ontology classes, including many musical sounds.Spain
GAPSClassical-guitar audio, scores and high-resolution MIDI alignments for transcription research.United Kingdom
GiantMIDI-PianoMachine-transcribed solo-piano MIDI derived from web recordings and IMSLP work metadata.Global
Groove MIDI DatasetHuman-performed expressive drum MIDI with aligned synthesized audio and performance metadata.United States
GuitarSetAcoustic-guitar recordings with hexaphonic audio and time-aligned note, chord, beat and technique annotations.United States
HEAR 2021A NeurIPS challenge and evaluation suite for general-purpose audio representations across speech, sound and music tasks.United States
HF2 Hardanger Fiddle DatasetPaired audio and MIDI dataset of Norwegian Hardanger fiddle material for automatic music transcription research.Norway
Kazakh Songs ASR DatasetGated non-commercial dataset of manually aligned Kazakh sung-vocal audio and text for ASR research.Kazakhstan / Kazakh language
KazEmoTTSKazakh emotional speech dataset for text-to-speech synthesis with six labeled emotions.Kazakhstan / Kazakh language
Lakh MIDI DatasetLarge deduplicated MIDI collection with matched and aligned subsets linked to the Million Song Dataset.United States
M6Multi-generator, multi-domain, multilingual, multicultural, multi-genre and multi-instrument music-detection databases.Global
MADBLarge music-aesthetics benchmark with multi-dimensional professional ratings and textual comments.Global
MAESTROPaired piano audio and high-precision MIDI from International Piano-e-Competition performances.United States