papersTODAY 04:00 UTC
Study Finds Clinical LLM Agents Give Inconsistent Orders Across Repeated Runs
A new arXiv paper examines how clinical LLM agents behave when given the same patient case multiple times. Although the agents often reach the same overall judgment, the tests, medications, and referrals they order can differ substantially between runs. The authors argue that evaluating these agents on a single run per task can hide this variability and misrepresent their reliability.