跳到正文

MoWAM: Explicit Future Motion Prediction for Efficient World Action Models

机器人 0 次浏览

原文MoWAM: Explicit Future Motion Prediction for Efficient World Action Models

作者:Jiayu Wang, Bin Zhu, Yue Yu, Jingjing Chen

来源:arXiv cs.RO(机器人)

正文

Computer Science > Robotics

arXiv:2609.20709v1 (cs)

[Submitted on 17 Sep 2026]

Title:MoWAM: Explicit Future Motion Prediction for Efficient World Action Models

Authors:Jiayu Wang, Bin Zhu, Yue Yu, Jingjing Chen

View a PDF of the paper titled MoWAM: Explicit Future Motion Prediction for Efficient World Action Models, by Jiayu Wang and 3 other authors

View PDF

HTML (experimental)

Abstract:World Action Models (WAMs) improve robot policy learning by incorporating future dynamics, yet explicitly generating future videos at inference introduces substantial computational overhead. Removing future generation improves efficiency, but leaves future dynamics only implicitly encoded in observation features, which can limit robustness under distribution shifts. We propose MoWAM, an efficient WAM that replaces future video generation with explicit future motion prediction. Instead of reconstructing the complete future scene, MoWAM models structured robot motion as a compact abstraction of the future, capturing how the robot is expected to evolve under the current scene and interaction constraints. A Mixture-of-Transformer architecture learns future visual dynamics during training while jointly predicting motion and action, allowing video generation to be removed entirely at inference while retaining an explicit representation of the future. The compact motion representation further enables efficient inference-time scaling by sampling multiple candidates of motion and action pairs and selecting among them with a motion-aware task-progress verifier. Experiments on LIBERO, LIBERO-Plus, and real-world manipulation tasks demonstrate that MoWAM achieves strong in-distribution performance, improved out-of-distribution robustness, and higher average real-world success than representative WAM baselines. In addition, performance improves as more candidates are explored, demonstrating that explicit future motion provides an effective and efficient basis for inference-time scaling.

Subjects:

Robotics (cs.RO)

Cite as:

arXiv:2609.20709 [cs.RO]

(or

arXiv:2609.20709v1 [cs.RO] for this version)

https://doi.org/10.48550/arXiv.2609.20709

Focus to learn more

arXiv-issued DOI via DataCite (pending registration)

主题

机器人


由「前沿雷达」于 2026-09-20 采集。正文取自原文页面,已保留出处链接。标题与正文版权归原作者所有。

评论

加载中…