The AI Race Gets Awkward as Western Labs Adopt DeepSeek Cache Optimizations
Summary
This opinion essay argues that the competitive dynamic between Chinese and Western AI labs has shifted from alleged model distillation toward adoption of openly shared optimization techniques. It focuses on DeepSeek’s reported progress in reducing the KV-cache footprint for long-context workloads such as coding. The author attributes this progression to methods including MLA, Compressed Sparse Attention, cross-layer cache reuse, a causal encoder-decoder design, and FP4 caching, and says the latest claimed result brings the cache to 890 bytes per token, roughly 437 times smaller than DeepSeek-V1 for the cited use case. Because KV cache occupies GPU memory during serving, the essay argues that the reduction can materially lower the cost of running long-context models. It further claims that Anthropic and OpenAI have released newer models using similar optimizations, naming Claude Opus 5.5 and GPT-6.1 Sol, and interprets their quieter launches as evidence that the technology was adopted without much public acknowledgment. According to the article, Opus 5.5 reduced cache-read pricing by 60% compared with Opus 5, while GPT-6.1 Sol reduced the comparable price by 80% against GPT-5.6 Sol’s late-July pricing. The author links the Chinese labs’ focus on efficiency to restrictions on access to advanced GPUs and argues that their openly shared techniques have helped Western labs improve inference margins. Claims about the companies’ motives, user reviews, model quality, and causal relationship between the optimization work and pricing are presented as the author’s interpretation rather than independently established evidence in the text.