AI Infrastructure Is Taking a Different Path From Classic Cloud
Summary
AI infrastructure is still immature, but its development is already diverging from classic cloud infrastructure. The article argues that AI infrastructure is built around machine-learning workloads: model training requires massive compute before deployment, while inference is the closest equivalent to running a web application. Most application and agent builders can use model APIs without understanding GPUs, except when they train, fine-tune, or self-host models. The market also reverses classic cloud’s concentration pattern: providers include hyperscalers and a growing group of GPU-focused neoclouds, while demand is concentrated among a small number of well-funded model labs and large inference providers. For inference, platform-as-a-service is expected to dominate, including among enterprises, because operating high-performance GPU systems is uncommon and difficult. Closed models such as GPT, Claude, and Gemini are API-only, and even many users of open-weight models rely on third-party inference providers. Customer requirements have not yet matured into the SLA, compliance, redundancy, and lock-in concerns associated with mature cloud markets because GPU availability remains the primary constraint. The article’s conclusion is that AI infrastructure is heading toward fragmented providers, concentrated consumers, PaaS-led consumption, and a long period in which access to compute matters more than refined procurement terms.