Feed

Trace inversion could weaken hidden-reasoning defenses against model distillation

postOriginal · 9 August 2026Revision 1
Preview image for Reasoning trace inversion and open-weight model distillation
Image from X

Jack Morris, a co-author of a 2026 paper on trace inversion, connects speculative rumors about Chinese labs extracting long-horizon reasoning from Claude Code and Codex to the recent strength of open-weight models. He is explicit that this account is unverified. The firmer result is his paper's finding that a model can infer useful synthetic reasoning traces from a frontier model's inputs and outputs, and that those traces improve student-model training over answers alone. Anthropic separately says it observed Moonshot trying to reconstruct Claude reasoning traces during a large-scale distillation campaign. Together, the evidence suggests that hiding chain-of-thought may not be a durable defense against capability extraction, which matters for both frontier-model protection and the future of open-weight models.