Back to News
RSS feedarxiv.org

Benchmark Separates Memory Retention from Retrieval in AI Agents

Summary

Persistent agents must decide what information to keep and what to retrieve when a query arrives, but memory evaluations can mix these two decisions. The authors build a streaming-recall benchmark that crosses retention and selection rules and tests every condition on the same 300 seeded episodes. With access held fixed, query-aware selection improves required-fact recall by 15.5 percentage points, with a 95% confidence interval of 12.8 to 18.2 points. A mixed comparison that also changes access reports a 68.7-point advantage, but 53.2 points of that difference are attributed to access rather than selection. Under bounded retention, query-aware, dense, and oracle selection all reach the retention ceiling, and all 319 observed failures in the bounded-recency condition result from eviction rather than ranking errors. Recall eventually falls to 0% when targets are far enough in the past. A repeat evaluation on SQuAD preserves the retention ceiling and shows that dense retrieval can outperform lexical retrieval on natural text. The authors conclude that bounded-memory evaluations should hold access constant and report retention and selection separately.