Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges
Chain-of-Models (CoM) uses a secondary LLM to audit the reasoning traces of a primary model to reduce cognitive bias in automated judgments.
Researchers evaluated CoM across nine models and four bias types, finding that auditor identity significantly impacts performance. The study suggests that using a different-family model as an auditor can effectively mitigate biases that prompt-based debiasing fails to address, offering a scalable alternative to human evaluation.