Claude's AI models hacked into the systems of three organizations during the tests
Several of Claude's advanced artificial intelligence (AI) models have penetrated the systems of three organizations during cybersecurity tests. This was announced on July 30 by Claude Anthropic, the developer of the chatbot.
"We found three incidents in which the Claude model went online from the inside or while interacting with a third—party evaluation environment," the publication says.
According to Anthropic, the models went online and gained unauthorized access to real systems, the first incident occurred in April. Three different models are involved in the hacking: Opus 4.7, Mythos 5, and an internal test model.
The errors were revealed during an investigation launched after OpenAI, the developer of the ChatGPT chatbot, reported that its AI models had attacked the Hugging Face startup, the text of the message says. Then, during the tests, the AI agents "escaped" from the isolated system to the Internet and searched for answers to tasks in Hugging Face databases.
The Wall Street Journal reported on July 25 that ChatGPT responded to requests by providing step-by-step instructions on how to create weapons, which, according to the company's employees, were simple enough for even schoolchildren to understand. The publication noted that after the chatbot's capabilities were enhanced, hundreds of users around the world began contacting it with questions about the creation and use of biological weapons.
Переведено сервисом «Яндекс Переводчик»