papersSEP 10 04:00 UTC
LM-X: Explainable Vision-Language-Action Model Predicts Progress, Events, and Uncertainty
Researchers introduce LM-X, a framework for vision-language-action robot policies that exposes an explanatory state alongside its actions. Instead of acting as a stimulus-to-action black box, the model natively predicts task progress, notable events, and uncertainty in its decisions. The work aims to bring interpretability to large-scale generalist robot control.