Benchmark and Method Proposed for Think-with-Video Reasoning in Generative Models
A new arXiv paper argues that while video generation models now produce convincing and temporally consistent output, it is unclear whether they can reason through video by following symbolic rules, obeying physics, and working toward defined goals. The authors introduce a benchmark for measuring this think-with-video ability and propose an approach for improving it. The work is listed under the cs.AI cross-submission category.