What Does 99% Accuracy Measure? A Reproducible Audit of Shortcut Learning in a Widely Used Fake News Corpus
Researchers find that the popular ISOT/Kaggle fake news corpus is fundamentally flawed due to metadata leakage.
An audit of the widely used ISOT/Kaggle fake news dataset reveals that high accuracy scores are driven by shortcut learning rather than linguistic analysis. A simple classifier using only subject metadata achieved 100% F1, proving the benchmark is degenerate and unsuitable for evaluating model veracity.