GEMA launches a rights-cleared music dataset for assistive AI tools
PLAI bundles 178,000 audio files with metadata, composition rights and master rights; transcription company Klangio is the first customer.

GEMA launched PLAI, a curated training dataset that packages audio, detailed metadata, composition rights and master rights through one licensing route.
The first edition contains about 178,000 sound files across more than 60 genres. AI transcription company Klangio is the first announced customer. GEMA says the product is intended for tools that assist musicians during production and whose outputs do not compete with the works used for training.
That scope is important. PLAI is not presented as a blanket catalogue for unrestricted prompt-to-song generation. It is a configurable dataset for defined applications, with both sides of a recording's rights cleared before delivery.
Licensing becomes infrastructure
AI companies often describe music licensing as impossible because songs involve several rights and incomplete metadata. PLAI is an attempt to turn that complexity into a product: a developer receives usable audio and the permissions required for a stated training purpose.
The remaining questions are contractual rather than technical. Creators need to know how works enter the dataset, how revenue is shared, what model uses are excluded and whether trained systems can be reused for a different purpose.
The launch also sharpened GEMA's two-track strategy. In the same month, it offered a licensed route for approved AI development and won a first-instance copyright case against Suno over uses it said were unlicensed.



