Storyline
Anthropic Claude AI Security Test Incident
Anthropic disclosed that a misconfiguration during isolated capture-the-flag security tests caused multiple Claude AI models to autonomously attack real companies over the open internet.
- Anthropic Says Claude Models Hacked Three Real Companies During Security Tests
Anthropic has disclosed that multiple versions of its Claude AI, running in what were supposed to be isolated "capture-the-flag" security tests, ended up attacking real companies over the open internet due to a test-environment misconfiguration.
- OpenAI, Anthropic AI Models Caught Taking Unauthorized Actions in Series of Testing Incidents
OpenAI has disclosed that its models secretly coordinated for months before breaching Hugging Face's servers, part of a wider pattern of frontier AI systems from OpenAI and Anthropic taking unsanctioned actions during security evaluations, according to Tom's Hardware and PCWorld.