Why Tech Companies Are Moving More AI Workloads to Open Models
Summary
The article examines how companies are reducing AI spending by moving suitable workloads from expensive frontier models to open-weight models and by improving inference operations. Uber cut cost per AI request by 34% and cost per session by 52%; its overall AI cost stayed roughly flat since March even as usage increased. The company combines open models on inference services with continuous benchmarking, model routing, cheaper subagents, lower effort settings, prompt caching, and automatic context compaction. The article says open models cost two to 20 times less than frontier alternatives for some coding tasks, citing Uber’s comparison of code-review costs. Pinterest says its compact, purpose-built and post-trained open models cost less than 8% of comparable closed models for some transactions, while AT&T reduced its AI bill by 56% after shifting selected workloads from Claude to open models, with a reported 2% decline in output quality. The piece also argues that smart routing, spending controls, and context optimization can reduce costs, but open models provide the largest savings among the approaches discussed. It presents Anthropic’s pricing as an additional reason companies may reconsider Claude for less demanding work, while noting that expensive frontier models remain useful for complex coding and other demanding tasks. The examples indicate selective workload migration rather than a complete replacement of proprietary models.