<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>统计机器学习 on 庞玉栋个人博客</title><link>https://pangyd.com/tags/%E7%BB%9F%E8%AE%A1%E6%9C%BA%E5%99%A8%E5%AD%A6%E4%B9%A0/</link><description>Recent content in 统计机器学习 on 庞玉栋个人博客</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Thu, 17 Sep 2026 23:10:18 +0800</lastBuildDate><atom:link href="https://pangyd.com/tags/%E7%BB%9F%E8%AE%A1%E6%9C%BA%E5%99%A8%E5%AD%A6%E4%B9%A0/index.xml" rel="self" type="application/rss+xml"/><item><title>Parallelism, critical windows, and separations among diffusion language models</title><link>https://pangyd.com/post/radar-376/</link><pubDate>Thu, 17 Sep 2026 23:10:18 +0800</pubDate><guid>https://pangyd.com/post/radar-376/</guid><description>&lt;blockquote>
&lt;p>&lt;strong>原文&lt;/strong>：&lt;a href="https://arxiv.org/abs/2609.20539v1">Parallelism, critical windows, and separations among diffusion language models&lt;/a>&lt;/p>
&lt;p>&lt;strong>作者&lt;/strong>：Sitan Chen, Liye Wang&lt;/p>
&lt;p>&lt;strong>来源&lt;/strong>：arXiv cs.LG（机器学习）&lt;/p>
&lt;/blockquote>
&lt;h2 id="正文">正文&lt;/h2>
&lt;p>Computer Science &amp;gt; Machine Learning&lt;/p>
&lt;p>arXiv:2609.20539v1 (cs)&lt;/p>
&lt;p>[Submitted on 17 Sep 2026]&lt;/p>
&lt;p>Title:Parallelism, critical windows, and separations among diffusion language models&lt;/p>
&lt;p>Authors:Sitan Chen, Liye Wang&lt;/p>
&lt;p>View a PDF of the paper titled Parallelism, critical windows, and separations among diffusion language models, by Sitan Chen and 1 other authors&lt;/p>
&lt;p>View PDF&lt;/p>
&lt;p>HTML (experimental)&lt;/p>
&lt;p>Abstract:A popular selling point of diffusion large language models (dLLMs) is their capacity for parallelism: the ability to generate sequences of text far more efficiently than autoregressive models, which require one forward pass per token. Yet among the many competing paradigms for dLLMs, from masked to uniform to Gaussian diffusion, principled understanding of how these different proposals compare in parallelism remains limited. In this work, we initiate a fine-grained comparison of the capacity for parallelism among these three leading approaches and prove the following:&lt;/p></description></item></channel></rss>