Back to News
RSS feedengineering.atspotify.com

How AI Changed Spotify’s Development: Quality Lessons at Higher Velocity

Summary

Spotify describes how AI-assisted development changed quality and reliability work across a platform serving 777 million monthly active users, about 100 million concurrent clients, 11–12 million backend requests per second, and nearly 3,000 production services. The company says AI was not the direct cause of the production incidents it reviewed, but the volume of change grew faster than some verification controls could adapt. In content processing, a scheduling bug, competing batch work, higher per-episode compute needs, and insufficient capacity delayed some episodes for hours; Spotify responded with end-to-end monitoring, scheduler fixes, lower-priority batch jobs, more capacity, and revised workload tiers. Its fleet-management system also supported more complex agentic changes, including a Java migration across backend services completed in three days, but an automated dependency upgrade still passed checks and failed in production, prompting stronger safeguards, rollback capacity, and working-hours scheduling. Industry-wide AI demand made CPU and GPU capacity less predictable, worsening the impact of regional failovers and leading Spotify to double reserved edge capacity while developing more gradual traffic spillover controls. In the mobile app, faster change exposed gaps in existing quality signals, so the company added broader metrics and longer-term trends. Spotify’s August merged changes more than doubled year over year from roughly 8,100 to 17,000, while quality and optimization work increased from 27% to 31% of the mix; the company reports no corresponding rise in rework rate, although code complexity and pull-request size are increasing. Spotify’s conclusion is that AI expanded its capacity to produce change, making verification, observability, rollback, resilience, and human engineering judgment the next constraints.