papersTODAY 04:00 UTC
Paper argues behavioral consistency is a distinct, measurable property of LLM agents
A revised arXiv preprint contends that current agent evaluations lean almost exclusively on outcome measures like success rate, which show whether an agent finishes a task but not how uniformly it behaves. The authors propose treating consistency of behavior across different tasks as its own measurable characteristic. Their work offers a way to evaluate agents beyond simple pass/fail scores.