papersTODAY 04:00 UTC
MemRiskBench Benchmark Targets Memory Risks in Long-Horizon LLM Agents
A new arXiv paper introduces MemRiskBench, an evaluation framework for long-horizon LLM agents that accumulate memory across sessions. It measures per-risk failure rates for issues such as stale facts, conflicting updates, cross-user data leakage, reuse of revoked memories, and decay of constraints, which aggregate scores tend to obscure. The work argues for trace-aware evaluation that preserves these distinct risk categories rather than collapsing them into a single number.