ArtifactBench Evaluates AI-Generated Music Detectors Under Distribution Shift
Summary
ArtifactBench is a lineage-aware evaluation suite for detectors of AI-generated music. The paper argues that aggregate benchmark scores can be misleading when training overlap, generator lineage, source provenance, and audio-transformation history are only partly known. Its protocol groups original recordings and derived variants by content identity, separates calibration data from final tests, records inference failures separately from classification errors, and reports source-level results with uncertainty. The authors evaluate several publicly available detectors under a version-pinned common protocol across generator families and versions, real-music domains, collection cohorts, and inference coverage. On a 562-track test intersection where all detectors produced results, ArtifactNet reached 0.982 AUROC and 0.918 balanced accuracy, while the public Deezer detector reached 0.761 and 0.776. SpecTTTra and CLAM both fell below 0.30 AUROC on the shifted cohort. Additional analysis shows that leakage controls, cohort availability, threshold policy, and detector-specific missingness can materially change measured performance and rankings. The paper also reports substantial shifts across generator families and real-music domains that aggregate scores conceal. It presents ArtifactBench as a benchmark and evaluation protocol rather than a new detector.