Four August papers test where AI-music evidence breaks

New studies examine musical sameness, broadcast detection, mixed AI stems and edited-audio false positives. Together they argue against one universal detector score.

First-page collage of four August 2026 AI-music research papers · arXiv authors, editorial crop by aimusic.events · 22 Aug 2026

Four preprints published in August examine different parts of the AI-music evidence problem. None is a universal verdict on generation quality or detection. Read together, they show why a result from clean, fully generated audio should not be carried unchanged into broadcast, edited or partly generated music.

Do generators make music more homogeneous?

The first study generated 100 tracks per system and genre with Suno and Lyria, then compared 72 music-information-retrieval features across four genres. It reports lower within-genre diversity for Lyria and weaker separation between genres for Suno in its test. The result is evidence about this prompt set, model set and feature representation — not proof that every output sounds the same. The paper is listed as forthcoming at AIES 2026.

A clean detector can fail on television audio

The BAMM paper assembled a 40-hour dataset from real television broadcasts. A detector that reached an F1 score of 0.992 on clean samples fell to 0.472 in the best broadcast setup; the clean-trained model scored 0.186 on the broadcast material. Speech, mixing, compression and other programme audio changed the task substantially.

The dataset covers Suno v3.5 rather than every generator. Its strongest conclusion is therefore operational: a benchmark on isolated files does not establish reliability after real broadcast processing.

Mixed stems are measurable — inside one controlled pipeline

A third paper constructed mixtures containing different proportions of generated stems and trained a regression model to estimate the share. It reports a mean absolute error of 0.076 and an R² of 0.85 within its setup.

The generated condition was produced through a controlled codec-reconstruction pipeline, not a broad sample of commercial generators and production chains. It is a promising experiment in moving beyond binary labels, but not yet a field-ready percentage meter.

Edited human audio is the hard negative

The fourth paper compares AI music with edited audio collected from YouTube. Its best system reached 0.811 balanced accuracy, while the edited-audio class had an F1 score of 0.720. That gap matters because mastering, restoration, time-stretching and heavy production can resemble artifacts a detector learned to associate with generation.

The practical reading

Ask what a detector was trained on, what transformations occurred after generation, whether the track is wholly or partly synthetic, and what false positives look like on edited human music. “Our model is 99% accurate” is not enough without the dataset and deployment condition.

These are recent preprints and workshop-stage results. They are useful for identifying failure modes, not for making irreversible authorship or enforcement decisions by themselves.

Original sources

Primary sourcearXiv
Primary sourcearXiv
Primary sourcearXiv
Primary sourcearXiv
Four August papers test where AI-music evidence breaks · AI Music Events