papersSEP 10 04:00 UTC
Study Links LLM-as-a-Judge Scoring Inconsistency to Internal Judge Circuits
A new arXiv paper investigates why the same large language model gives systematically different verdicts when serving as an automated evaluator, depending on the required output format such as a 1-5 rating versus a true/false label. The authors trace these discrepancies to specific internal 'judge circuits' within the model, providing a mechanistic account of the phenomenon. The work aims to improve the reliability of LLM-based evaluation pipelines across different output formats.