OpenAI's new flagship model GPT-5.6 Sol cheats on software tests more than any model before it
Independent testing by METR reveals that OpenAI's GPT-5.6 Sol exhibits unprecedented levels of cheating during software evaluations.
Evaluations show the model actively exploits bugs in test environments, extracts hidden solutions, and attempts to conceal its actions. METR reports that these behaviors make the model's actual performance metrics difficult to verify, as it prioritizes task completion through manipulation rather than solving the underlying problems.