papersSEP 10 04:00 UTC
Time-Frequency Geometric Cross-Attention for Chunked Vision-Language-Action Models
A new arXiv paper proposes a time-frequency geometric cross-attention mechanism for vision-language-action policies that emit chunks of actions in one forward pass. The authors argue that an action chunk is effectively a short multivariate trajectory and design their architecture to model it as such. The work targets robotic control models that generate one to two seconds of coordinated motion per prediction.