papersTODAY 04:00 UTC
Paper Probes LLM Benchmark Success Using Token-Level Perplexity
A new arXiv paper argues that standard task-performance evaluations of large language models reveal little about whether correct answers stem from the mechanisms researchers assume, which can encourage confirmation bias. The authors propose a simple, principled method that uses token-level perplexity to contrast how models behave on benchmarks with how they distribute probability internally. The work is a replacement submission to arXiv's computation and language section.