papersSEP 10 04:00 UTC
RAP benchmark probes how LLM research agents track shifting scientific attention
Researchers introduced RAP, a task for measuring whether large language models serving as research agents can follow changes in scholarly attention, which has been hard to assess because reviews and proposed ideas lack verifiable outcomes. The findings indicate that these models acquire evidence in ways that are biased toward the specific target they are given.