Back to News
RSS feedepoch.ai

How Many AI Agents Could Run on Chips Shipped Through 2027?

Summary

Epoch AI estimates how many concurrent agentic workloads could run on high-bandwidth memory shipped from 2025 through 2027. Under its central assumptions, the hardware could support about 30–170 million concurrent frontier-model agents once fully deployed and allocated to these workloads, equivalent to the weekly working hours of roughly 140–720 million full-time employees. More efficient open models could raise the estimate substantially: applying DeepSeek V4 Pro serving benchmarks produces approximately 1.9 billion concurrent agents in one scenario. The analysis combines HBM shipment estimates with serving benchmarks for open models and API-equivalent spending for closed models. Its reference case assumes $30 per active agent-hour, $5 per GB300-hour, a 5–10x API-revenue-to-serving-cost ratio, and that HBM4/4E supports twice the agents per unit of memory as HBM3E; the tested hardware uplift ranges from one to four times. Agent-hour spending in the analyzed traces varied by model and harness, from about $15.50 for GPT-5.6 Sol and $18.19 for GPT-5.5 Codex sessions to $24.34 for Opus 4.8 and $50.16 for Fable 5 Claude Code sessions. With only 20% of the central capacity effectively used for revenue-generating inference, Epoch estimates $2.6–5.3 trillion in annual API-equivalent spending after 2027 shipments are deployed. That capacity could exceed demand if agent adoption, delegated workloads, and willingness to pay do not grow quickly enough, although deployment delays may give demand more time to catch up.