EngramRAG Builds Usage-Adaptive Memory for Multi-Hop LLM Agents
Summary
EngramRAG is an adaptive memory architecture for autonomous LLM agents operating across multiple sessions. The paper argues that conventional memory systems suffer from three problems: they cannot reliably traverse multi-hop relationships, may discard stable persona information through time-based decay, and use static graph structures that do not reflect actual usage. Inspired by Complementary Learning Systems, EngramRAG separates a low-latency Waking State retrieval reflex from an asynchronous Dreaming State consolidation cycle. Its main mechanisms are Usage-Modulated Personalized PageRank, which uses Hebbian-style usage updates to promote persistent entities into high-centrality hubs; Consolidation-Activated Topology Decay, which adjusts retention half-life according to a node’s structural importance and applies a cold-start grace period of at least four observations; a directed SUPERSEDES graph that filters obsolete facts after mutations; and dynamic Reciprocal Rank Fusion across dense-vector, BM25, and graph-based retrieval. On all 1,982 question-answer pairs from 10 long-term LoCoMo conversations, EngramRAG raises Recall@5 from 38.29% for dense-vector RAG to 53.21%, a 38.9% relative improvement with p < 0.001. Its MRR is 0.4203 versus 0.2937, and it also exceeds Okapi BM25 at 48.66% and isolated static graph retrieval at 8.50%. For temporal reasoning, it reaches 62.33% Recall@5, 16.67 percentage points above dense vectors. In controlled fact-mutation tests, SUPERSEDES reduces split-brain hallucinations from 70.0% to 0.0%. A 90-day simulation reports 100.0% scaffolding retention while maintaining a 26.21 ms interactive retrieval reflex. These results are reported for the proposed benchmark and simulations; the abstract does not establish performance in other deployment settings.