🤖 AI 资讯

· ·
← 返回列表

What Does Chain-of-Thought Entropy Measure? A Channel Audit of Scaffolding, Routing, and Content

arXiv cs.LG2026-09-23 04:00:00扩散模型,微调蒸馏,招聘HR,论文原文 ↗

arXiv:2609.25039v1 Announce Type: cross

Abstract: Entropy over chain-of-thought tokens decides which tokens receive the policy gradient, which get pruned, and whether a run has collapsed, yet each such statistic reads a next-token distribution mixing three choices: whether to emit connective scaffolding, which connective, and what the substantive continuation should be. Designating a scaffold vocabulary subset separates the three, exactly, for entropy, Kullback--Leibler divergence, and the first-order entropy velocity of a softmax policy. We prove the raw and content conventions disagree about which position is the larger fork on an explicit open region, and bound answer diversity by the content channel plus a leakage term a measured witness certifies. Across twenty-three configurations the scaffold side carries up to 41% of the raw high-entropy set; on a matched-tokenizer ladder, coupling changes only at the math-corpus step while the scaffold's entropy share keeps growing through distillation; a closed-form forecast from one channel correlation tracks selection retention over a 54-point range to five points, unfitted. On compression, the content convention beats raw surprisal in every cell; an answer-leakage audit then corrects our own headline control: re-fed chains earn a quarter to a half of their accuracy from restated answers, and once stripped, no token scorer beats a random contiguous block.