papersTODAY 04:00 UTC
Audit questions whether test-time scaling pays off for video world models
A new arXiv paper argues that adding inference-time compute helps video world models only when the extra samples are actually better and can be reliably picked out. The authors separate the benefit of a larger candidate pool from the ability to select the best one, an effect they call sampling headroom versus selection gain. They propose auditing test-time scaling along this distinction to judge whether the added compute is worth its cost.