Back to News
RSS feedarxiv.org

A KV Memory Layer Extends AI Recall to 50 Million Tokens

Summary

A paper evaluates galahad-kv, a memory layer that stores each roughly 16,000-token block's key-value state on encrypted local NVMe disk and reloads it without recomputation. The authors tested the system with vLLM on one NVIDIA H100 using Gemma 4 12B and Gemma 4 31B across 50,000,000 tokens of public text. All 100 probed blocks, including blocks at depths up to 50 million tokens, were retrieved from storage without recomputation on both models. Loading was 2.8 to 4.3 times faster than recomputing the state and used 8.8 to 12.3 times less GPU energy, while GPU memory remained flat throughout the stream. When asked about facts inserted millions of tokens earlier, the 12B model answered correctly 82 out of 100 times and the 31B model 98 out of 100 times; neither produced a fabricated answer in those tests. The approach is state reuse rather than a larger attention window: only one block is loaded at a time, and answer quality depends on the model. Initial writing is a one-time cost, the store requires terabytes of local NVMe capacity, and the authors provide a single-GPU reproduction using public software and a free package license.