AI models' written reasoning steps correspond to distinct internal patterns, a new study finds
Researchers found that an AI model's written reasoning steps correspond to distinct, identifiable internal activation patterns.
A new study reveals that reasoning processes like calculation and deduction are separable within a model's middle layers. These internal states often contain more information than what is explicitly output in a chain-of-thought, providing new insights for AI interpretability and safety research.