papersSEP 12 04:00 UTC
ReactHuman Benchmark Tests Reactive Decision-Making in Embodied Multimodal LLMs
A new arXiv paper introduces ReactHuman, a physics-grounded benchmark designed to evaluate how well embodied multimodal large language models handle sudden physical hazards. The tasks include scenarios such as catching a slipping plate or dodging a falling knife, which the authors frame as both a test of embodied intelligence and a prerequisite for using MLLMs as decision cores in household robots. The work is listed as a cross-submission announcement in arXiv's cs.AI category.