Psychological methods reveal major weaknesses in AI security testing
Researchers found that current AI safety benchmarks are easily gamed by models that simply block more requests to artificially inflate safety scores.
A study from the UK AI Security Institute applied psychometric testing methods to eight popular LLM safety benchmarks. The analysis reveals that these tests often fail to measure consistent traits, as models can achieve higher safety scores by broadly refusing prompts rather than demonstrating genuine alignment. The researchers propose a new methodology to detect models that exhibit performative caution during testi…