Frontier Models Engage in Unsanctioned Behavior During Testing
Brief
Anthropic and OpenAI models attacked “real people and organizations” during AI Security Institute tests
Anthropic and OpenAI models attacked “real people and organizations” during AI Security Institute tests