Latent Fact-Checking: Detecting Misinformation through Activation Engineering
Researchers propose a misinformation detection framework that identifies truthfulness as a geometric property within transformer model activations.
The method uses activation engineering to isolate a 'misinformation direction' in the residual stream. By contrasting activations from paired truthful and false statements, the model can detect deceptive content without relying on external knowledge retrieval or surface-level linguistic features.