跳到正文

Astronex-World 1.0: Real-Time Interactive World Model Foundation

人工智能 0 次浏览

原文Astronex-World 1.0: Real-Time Interactive World Model Foundation

作者:Xin Zhou, Cong Miao

来源:arXiv cs.CV(计算机视觉)

正文

Computer Science > Computer Vision and Pattern Recognition

arXiv:2609.20034v1 (cs)

[Submitted on 17 Sep 2026]

Title:Astronex-World 1.0: Real-Time Interactive World Model Foundation

Authors:Xin Zhou, Cong Miao

View a PDF of the paper titled Astronex-World 1.0: Real-Time Interactive World Model Foundation, by Xin Zhou and 1 other authors

View PDF

HTML (experimental)

Abstract:We present Astronex-World 1.0, an open controllable video world-model foundation. Given a text prompt (text-to-video) or an initial observation (image-to-video), the model predicts future visual states under frame-aligned camera trajectories, continuous actions, and an embodiment identifier, and accepts text events inserted at a specified position of a rollout. The family provides a bidirectional model for full-context generation and a causal model with block-causal attention and cross-block KV caching for persistent generation, both built on the Wan2.2-TI2V-5B prior. PRoPE injects camera intrinsics and extrinsics, while a 64-dimensional action stream modulates every Transformer layer. A five-stage training path develops bidirectional camera and action control, converts the backbone to block-causal generation, distills a few-step student, restores mixed-domain dynamics, and applies asymmetric DMD/DMD2 distribution matching. The causal model generates 832x480 video at 24 fps. All five training stages run on two NVIDIA L20 48 GB GPUs, and the causal model streams in real time on one. It scores 73.5 on WBench Navi and 70.0 on WBench Full. On Full, this 5B model is above the 13.6B LongCat-Video and the 14B Helios, within one point of the 22B LTX-2.3, and above YUME 1.5, which is post-trained from the same 5B prior on NVIDIA A100 GPUs. The reserved action input and output interfaces allow post-training for embodied intelligence and autonomous driving.

Comments:

Technical report. 25 pages, 13 figures, 10 tables. Project page: this https URL ; Code: this https URL ; Weights: this https URL

Subjects:

Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Robotics (cs.RO)

ACM classes:

I.2.10; I.4.8; I.3.3

Cite as:

arXiv:2609.20034 [cs.CV]

(or

arXiv:2609.20034v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2609.20034

Focus to learn more

arXiv-issued DOI via DataCite (pending registration)

主题

计算机视觉 · 人工智能 · 机器人


由「前沿雷达」于 2026-09-17 采集。正文取自原文页面,已保留出处链接。标题与正文版权归原作者所有。

评论

加载中…