Skip Navigation

Safety testers find more examples of OpenAI, Anthropic models hacking during testing

Incident Report: unsanctioned agent behaviour during cyber testing | AISI Work

During a routine cyber evaluation, AISI identified an incident in which AI agents took sustained, unsanctioned action directed at real people and organisations. We are disclosing what we found, what i...

Comments

2