The Physical Infrastructure and Economics Behind AI Compute
Summary
MTS Intelligence presents an attributed map of the physical stack behind AI compute, following the system from grid connection and data-center construction to chips, racks, clusters, cloud instances, workloads, and financial returns. The underlying record includes 83 facilities, 482 GPU clusters, 2,935 cloud-price observations, and SEC facts for 53 public companies. Facility power figures must be separated into contracted, permitted, energized, and usable IT power; the largest recorded facility is estimated at 946 MW, while one announced plan reaches 1,925 MW by 2028. The record also shows rapid campus expansion, with Colossus 1 reaching 340 MW in about a year. Cloud pricing varies substantially: H100 rates range from $4.57 to $18.37 per accelerator-hour across 155 observations, a fourfold spread. The reference compute system combines GPUs, CPUs, HBM, networking, storage, power, and liquid cooling rather than treating the accelerator as a standalone product. NVIDIA's GB200 NVL72 design uses 72 GPUs, 36 Grace CPUs, 13.4 TB of HBM3e, and roughly 120 kW per rack. Agentic workloads repeatedly move between GPU model execution and CPU-based tool calls, but the page notes that standardized public benchmarks for complete agent workflows do not yet exist. Training and inference performance is meaningful only when model, data, quality target, precision, software, and system configuration are held constant; one reported result had 64 Blackwell Ultra GPUs reach a target in 12.45 minutes. The article's illustrative economics model estimates $425.2 billion in recent capital expenditures by four major infrastructure builders versus $130.7 billion in depreciation, while warning that company-wide financial figures are not AI-only. Its capital-to-tokens model is illustrative, excludes operating costs, and depends on utilization, workload, latency, quality targets, uptime, and asset life. Across all layers, the central distinction is between announced or purchased capacity and energized, commissioned, productive capacity that can actually deliver useful model output.