KnowBench proposes effort-reduction benchmark for clinical AI evaluation
A new arXiv preprint introduces KnowBench, a benchmark that assesses clinical AI systems by how much work they save clinicians instead of how closely their outputs match reference texts or expert rubrics. The authors argue that existing evaluation methods were built for research settings and measure resemblance to an artifact rather than reduction of a real-world burden. The paper frames deployment-grounded effort reduction as a unified metric for clinical AI.