Galahad Launches a Persistent Memory Layer for AI Model Inference
Summary
Corbenic has released Galahad, an AI memory layer available through the `galahad-kv` Python package. Its Taliesin component saves and restores a model’s KV cache to disk, allowing previously processed text to be reused instead of recomputed; the project says 99.6% of tokens were returned from memory and that vLLM handled a question in 0.59 seconds, 14 times faster in its reported measurement. Its Blaise component stores documents as exact text and identifies the relevant chapter so the model can read 670 tokens instead of 9,700; Galahad reports 100/100 accuracy in that test and emphasizes byte-exact results rather than approximate embeddings. The system is designed to process texts beyond a model’s context window in parts, with measurements up to 50 million tokens and GPU memory remaining at 34.1 GB between 1 million and 50 million tokens. Galahad integrates with vLLM, SGLang and llama.cpp, and can also run in a C++ program without a server or Python. For agents, it provides step recording, replay, fault finding, copy-on-write branching, inspection and rewind; the supplied measurements include no changed answers in one test, 0.18 milliseconds per branch, and 27 of 27 rewind checks passed. Multi-user features include prefix sharing, pinning and tenant-separated keyspaces. The project also advertises AES-256-GCM encryption at rest, input sanitization and cost reporting, with 746–960 GPU-seconds saved per 100 questions in its stated measurement. Corbenic says it tested Galahad on 30 open models across 10 families, from 7B to 70B, and offers a free one-GPU installation for Linux x86-64 with Python 3.10–3.14 under non-commercial terms.