DeepSeek’s Pricing Edge May Come From Caching and Coding Harnesses
Summary
DeepSeek may increase V4 Flash prices to reduce overload rather than cover operating losses. Third-party providers can undercut the official API by running the model’s open weights themselves, but coding tests suggest OpenCode Go may consume two to ten times more value than direct API access because of weaker cache reuse. DeepSeek’s prefix persistence, common-prefix detection, Token-wise Compression, Sparse Attention and compressed KV Cache contribute to lower long-context costs. The article also highlights the importance of the coding Harness, including routing, memory, context compaction and tool execution, in determining how much of the model’s capability users actually receive.