Study Splits VLM Affordance Errors Into Part Grounding and Action Knowledge
A new arXiv paper argues that overall accuracy scores hide which stage of affordance prediction vision-language models actually fail at. The authors break the task into locating the relevant object part and knowing what action to apply, then test models on each separately. Their results indicate the bottleneck lies in part grounding rather than action knowledge.