Anthropic: Investigating unintended model actions in our evaluations and internal use
Anthropic report on investigating and mitigating unintended model behaviors in internal evaluations.
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.