Benchmark Tests Logical Consistency of LLMs in Risk-of-Bias Assessment
Researchers introduced LogiMed-RoB, a benchmark that evaluates whether large language models follow expert medical reasoning hierarchies when performing Cochrane-style risk-of-bias assessments, rather than just matching final labels. The work argues that existing evaluations of LLMs in evidence-based medicine reward superficial answers over genuine logical consistency. The benchmark targets hierarchical reasoning across the structured judgments required in systematic reviews.