papersTODAY 04:00 UTC
Study Finds LLM Judges Underuse Non-Directional Verdicts Allowed by Task Contracts
A new arXiv paper examines how large language models act as judges in evidence-based fact verification, converting supporting material into final verdicts. The authors report that even when task instructions explicitly permit non-directional outcomes such as "Conflicting" or "Not Enough Evidence," models tend to favor directional verdicts instead. The work suggests a mismatch between stated judging criteria and the labels models actually produce.