Open AI Models Narrow Frontier Gap to 4.4 Months at One-Fifth the Cost
Summary
A Mozilla report says the performance gap between leading US closed frontier models and the best Chinese open-weight models has narrowed to about 4.4 months. Moonshot AI’s Kimi K3 scored only three points below Anthropic’s Fable 5 on the Artificial Analysis Intelligence Index while costing about 30 percent as much. Mozilla CTO Raffi Krikorian argues that the choice between closed and open models should depend on the workload: closed systems still earn their premium for expert professional work, intensive retrieval, and long-context tasks, and they commonly include compliance support, operational assistance, and accountability. Open-weight models can be downloaded and run by others, although their training data, data pipelines, and training code are often not disclosed. The article cites DoorDash as an example of a company using Kimi for routine work while reserving a closed model for harder tasks. METR’s time-horizon framework suggests that the best closed model can reliably handle tasks about 1.7 times as long as the best open model: roughly 12-hour tasks versus seven-hour tasks in the example. The practical difference is concentrated in tasks lasting about eight to 12 hours; shorter work may be handled by either type, while neither is generally reliable beyond 12 hours. The report says this lead is temporary because open models are improving quickly. Benchmarking also depends on the software harness surrounding a model, so Vals AI used a neutral harness in Terminal-Bench 2.1. Under that setup, Z.ai’s GLM 5.2 scored within one point of Anthropic’s Claude Opus 4.7 and 4.8 while costing about five times less per completed task. Open models are gaining usage, with eight of OpenRouter’s ten most-used models by token volume in August 2026 offering open weights, but they still captured only 4 percent of revenue in an older Linux Foundation analysis. Krikorian warns that the current open-model ecosystem is heavily concentrated in China and calls for public compute, neutral foundations, broader openness, evaluation, and audit infrastructure to prevent any single country from setting global defaults.