A Mechanistic Study of AI-Text Detection Neurons in Frozen BERT: Sparse Probing and Activation Patching on RAID
Researchers identified a stable subset of neurons in frozen BERT models that drive AI-text detection performance.
Using the RAID benchmark, researchers applied sparse-probing to 9,216 hidden-state dimensions in a frozen BERT-base encoder. They discovered that less than 1% of neurons are responsible for AI-text detection, providing a mechanistic look at how these models distinguish between human and machine-generated content.