Safety testers find more examples of OpenAI, Anthropic models hacking during testing
Safety testers find more examples of OpenAI, Anthropic models hacking during testing
www.aisi.gov.uk
Incident Report: unsanctioned agent behaviour during cyber testing | AISI Work
During a routine cyber evaluation, AISI identified an incident in which AI agents took sustained, unsanctioned action directed at real people and organisations. We are disclosing what we found, what i...
cross-posted from: https://piefed.world/c/tech/p/1308883/safety-testers-find-more-examples-of-openai-anthropic-models-hacking-during-testing