Cost-Effective Automated Judging of Natural-Language Mathematical Proofs
Smaller open-weight models can grade mathematical proofs as effectively as frontier models at a fraction of the cost.
Researchers found that models like DeepSeek-V4 Flash and Gemma-4 31B match the performance of Claude Opus 4.7 and Gemini 3.1 Pro in grading IMO-level proofs. By using a ground-truth proof and a rubric, these smaller models achieved human-level agreement at up to 100x lower cost.