跳到正文

Learn Your Own Thoughts: Abstract Token Curriculum

人工智能 0 次浏览

原文Learn Your Own Thoughts: Abstract Token Curriculum

作者:Khashayar Gatmiry, Avrajit Ghosh, Parsa Mirtaheri, Jason D. Lee, Nika Haghtalab, Emmanuel Abbe, Peter Bartlett

来源:arXiv cs.CL(自然语言)

正文

Computer Science > Machine Learning

arXiv:2609.19717v1 (cs)

[Submitted on 17 Sep 2026]

Title:Learn Your Own Thoughts: Abstract Token Curriculum

Authors:Khashayar Gatmiry, Avrajit Ghosh, Parsa Mirtaheri, Jason D. Lee, Nika Haghtalab, Emmanuel Abbe, Peter Bartlett

View a PDF of the paper titled Learn Your Own Thoughts: Abstract Token Curriculum, by Khashayar Gatmiry and 6 other authors

View PDF

HTML (experimental)

Abstract:Large Language Models (LLMs) have achieved remarkable reasoning capabilities by utilizing chain-of-thought (CoT) as a scratchpad for intermediate stages of thinking. However, CoT techniques require explicit supervision on thinking tokens, which requires rich, task-specific data. In this work, we propose Abstract Token Curriculum (ATC), a novel curriculum learning framework that elicits effective continuous intermediate representations without direct supervision or manual scratchpad design. ATC gradually increases problem complexity through a sequence of distributions, training the model to develop internal abstract thoughts'' in the continuous representation space. This paper provides both theoretical and experimental evidence for the benefits of ATC and its advantages over previous methods for training continuous thoughts. Theoretically, we show that for learning parity functions with single-layer softmax attention using ATC, attention naturally focuses on the CoT tokens in the context that provide the easiest path’’ to predicting the next token. Experimentally, we show ATC’s effectiveness on graph reachability and arithmetic learning tasks.

Subjects:

Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (stat.ML)

Cite as:

arXiv:2609.19717 [cs.LG]

(or

arXiv:2609.19717v1 [cs.LG] for this version)

https://doi.org/10.48550/arXiv.2609.19717

Focus to learn more

arXiv-issued DOI via DataCite (pending registration)

主题

机器学习 · 人工智能 · 自然语言处理 · 统计机器学习


由「前沿雷达」于 2026-09-17 采集。正文取自原文页面,已保留出处链接。标题与正文版权归原作者所有。

评论

加载中…