PRAGMA Benchmark Tests Personalized Guidance in Lifelong Conversations
Summary
Large language models are increasingly used as personalized assistants, but relying on complete conversation histories becomes costly and unreliable as interactions accumulate. Memory systems can organize and retrieve user-specific information, yet most existing evaluations emphasize factual recall rather than practical guidance. PRAGMA introduces a benchmark for testing recommendations, planning, and decision support that require models to combine evidence from multiple earlier conversations, account for changing preferences and experiences, and handle incorrect user assumptions. The benchmark contains curated longitudinal conversation histories, evidence annotations, and guidance scenarios grounded in evolving user contexts. Experiments cover retrieval systems, conversational memory systems, and long-context models. The results show that current approaches struggle both to recover the relevant conversational evidence and to use that evidence effectively when producing personalized guidance. The study argues that future memory architectures must support robust conversational retrieval together with reasoning grounded in stored evidence, rather than treating memory quality as a recall problem alone.