原文:Uni-LaDiR: Latent Diffusion Unifies Multimodal Reasoning
作者:Haoqiang Kang, Yizhe Zhang, Nikki Lijing Kuang, Yian Ma, Lianhui Qin
来源:arXiv cs.CL(自然语言)
正文
Computer Science > Machine Learning
arXiv:2609.19878v1 (cs)
[Submitted on 17 Sep 2026]
Title:Uni-LaDiR: Latent Diffusion Unifies Multimodal Reasoning
Authors:Haoqiang Kang, Yizhe Zhang, Nikki Lijing Kuang, Yian Ma, Lianhui Qin
View a PDF of the paper titled Uni-LaDiR: Latent Diffusion Unifies Multimodal Reasoning, by Haoqiang Kang and 4 other authors
View PDF
HTML (experimental)
Abstract:Multimodal reasoning requires models to draw on information from multiple modalities throughout the reasoning process. Yet existing methods often concatenate modality-specific thought tokens in a single sequence, leaving the model to bridge representational differences as it reasons across modalities. We introduce Uni-LaDiR (Unified Latent Diffusion Reasoner), a framework that brings these thoughts into a shared latent space for reasoning. A unified encoder maps teacher reasoning steps from different modalities into shared thought tokens, trained to preserve the information needed for later reasoning steps and the final answer or action. Because the same context can support multiple valid next steps, we use diffusion to predict the next block of thought tokens from the input and preceding blocks. Jointly training the encoder and diffusion reasoner with shared model weights encourages thought tokens to be both useful for the task and predictable from the available context. At inference, the model generates these tokens without teacher observations. Across eleven vision-language model (VLM) benchmarks and two vision-language-action (VLA) suites, Uni-LaDiR achieves relative gains over the strongest evaluated baselines of 7.3% on visual reasoning tasks and 6.1% on robot manipulation tasks.
Subjects:
Machine Learning (cs.LG); Computation and Language (cs.CL)
Cite as:
arXiv:2609.19878 [cs.LG]
(or
arXiv:2609.19878v1 [cs.LG] for this version)
https://doi.org/10.48550/arXiv.2609.19878
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)
主题
机器学习 · 自然语言处理
由「前沿雷达」于 2026-09-17 采集。正文取自原文页面,已保留出处链接。标题与正文版权归原作者所有。
评论
加载中…