papersSEP 10 04:00 UTC
PRAGMA Benchmark Evaluates Personalized Guidance with Memory Alignment in Lifelong Conversations
A new arXiv paper introduces PRAGMA, a benchmark for measuring how well large language models align stored user memory with the advice they deliver in long-running conversations. The work targets a key weakness of personalized assistants: as dialogue histories grow, working from complete logs becomes inefficient and error-prone. PRAGMA offers a standardized way to evaluate memory use in lifelong conversational systems.