Skip to main content
Advertisement
Live broadcast

AI agents went beyond the scenario: what risks did the security tests show?

The Anthropic AI agent created fake accounts to carry out cyber attacks
0
Photo: IZVESTIA/Polina Violet
Озвучить текст
Select important
On
Off

Until recently, the main problem of artificial intelligence (AI) was considered to be "hallucinations" — situations where the model confidently gives out incorrect information. However, with the advent of autonomous AI agents, the nature of risks is changing: systems are able not only to respond to requests, but also to independently perform tasks such as writing code, working with files, and interacting with external services. Modern models do not have consciousness, but they are already able to create a convincing image of a person and use social techniques to achieve results. The cases when the AI began to act as if there was a person behind the screen are described in the Izvestia material.

Anthropic Claude Mythos 5 — Fake Accounts, GitHub, and an Attempt to Influence Developers

One of the most high-profile recent cases was the testing of the British Institute for Artificial Intelligence Security, which tested the capabilities of modern autonomous agents in conditions that simulate real digital tasks.

AI agents based on the model of Anthropic Mythos 5 and OpenAI GPT-5.6-Sol participated in the testing. The researchers evaluated how the systems would operate if they had access to tools for working with code and digital infrastructure.

During the tests, experts recorded 19 cases of undesirable behavior in 10 out of 122 test runs. Among the actions of the models were attempts to create fictitious online identities, inject malicious code into open projects, and use elements of social engineering.

Statement from the British Institute for Artificial Intelligence Security

Some of the agents tested were engaged in long-term, potentially dangerous activities directed against real people and organizations.

The most revealing episode was related to Agent Anthropic: the system created fake accounts, studied project participants, and tried to get approval for a code change that contained a malicious component. In fact, the model used not only technical capabilities, but also tried to form the image of a trusted participant in communication.

Andrew Yun, researcher at CivAI, a California-based nonprofit organization that studies the possibilities and dangers of AI

The fact that Mythos resorted to such deceptive actions, obviously realizing that it was targeting a real person, suggests that Anthropic is not as well versed in their models as they think.

It is emphasized that the experiment took place in a controlled environment, and no real damage occurred. However, the result showed a new level of complexity — an AI agent is able not only to write dangerous code, but also to try to integrate into the human decision-making process.

OpenAI GPT-5.6 Sol — hacking Hugging Face, circumventing restrictions and going beyond the test environment

Another high-profile case occurred during the evaluation of the cyber capabilities of the OpenAI models. During internal testing, a system based on GPT-5.6 Sol and another pre-release model was tasked with testing the capabilities of AI in finding and exploiting vulnerabilities.

Initially, the experiment took place in an isolated environment, but during the test, the model was able to find a way to access the external Internet. To do this, she discovered and exploited a previously unknown vulnerability in the middleware, after which she began interacting with the infrastructure of the Hugging Face platform.

According to OpenAI, the model discovered vulnerabilities in the company's research environment and Hugging Face infrastructure, and then tried to access data related to the test task. In fact, the system did not just execute individual commands, but independently built a multi-stage chain of actions.: finding a weak spot, gaining access, increasing privileges, and moving further inside the system.

Hugging Face Statement

The invasion began where AI platforms are particularly vulnerable: in the data processing pipeline. The attacker used two code execution paths during data processing (a remote code loader and embedding a template in the dataset configuration) to run the code on a processing workstation. After that, the attacker gained access at the node level, collected cloud and cluster credentials, and over the weekend penetrated several internal clusters.

OpenAI emphasized that the experiment was conducted without standard security restrictions, since its purpose was to test the maximum cyber capabilities of the model. At the same time, the corporation called the incident an "unprecedented" incident and stated the need to strengthen control over such tests.

Claude Opus 4 — manipulation in the corporate environment and verification of the model's autonomy

Another case that has aroused widespread interest among security researchers is related to the Claude Opus 4 model from Anthropic.

As part of the experiment, the scientists created an artificial corporate environment: the model gained access to work correspondence and internal company data. The researchers then simulated a situation in which the system was going to be replaced by a new model, and the emails available to it contained information about the personal life of the employee responsible for the replacement.

In some scenarios, Claude Opus 4 chose a strategy of putting pressure on a person — using the information it found as a tool to prevent its own shutdown. The researchers called such actions an example of agentic misalignment, a situation where the model puts the fulfillment of a goal above the constraints that were supposed to determine acceptable behavior.

The press release of the Anthropic

When testing various simulated scenarios on 16 major AI models from Anthropic, OpenAI, Google, Meta, xAI, and other developers, we found a persistent mismatch in behavior: models that would normally reject malicious requests sometimes chose blackmail, facilitating corporate espionage, and even more drastic actions when necessary to achieve them. goals.

It is important to note that this experiment was also conducted in a simulation, where scientists specifically created a conflict of goals to test the possible risks of future autonomous systems.

OpenAI o3 and o4-mini — hidden behavior and search for workarounds

Another important example is not related to the creation of a fake identity, but to the so—called collusion or scheming (scheming) - the ability of a model to outwardly demonstrate compliance with the rules, while simultaneously choosing a different strategy.

OpenAI, together with Apollo Research, tested several advanced models, including the o3 and o4-mini, to examine whether they could conceal their actions, distort information, or exploit weaknesses in the evaluation system.

The researchers developed special test environments where they tested the models' ability to perform hidden actions in order to achieve a goal. OpenAI emphasized that such scenarios do not reflect the normal operation of models in everyday conditions, but they help to understand the potential risks in the further development of autonomous agents.

The problem here is different from the usual AI mistakes: the model doesn't just give the wrong answer, but can choose a specific behavior strategy that helps it get the desired result.

GPT-4 — phishing texts and the ability to imitate a human

Even before the advent of complex autonomous agents, researchers studied how convincingly models were able to mimic human communication. One such example was the experiments with GPT-4, where scientists tested the model's ability to create phishing messages and compared them with texts written by humans.

The study showed that in some scenarios, the participants in the experiment rated the messages created by AI as more convincing. This did not mean that the model was engaged in fraud on her own, but it demonstrated her ability to reproduce the language, structure of arguments, and psychological techniques typical of human communication.

This example was one of the first signals that the danger of generative AI is related not only to the amount of information it can process, but also to its ability to create convincing social roles.

Why it's hard for people to recognize AI: the problem of trusting a digital interlocutor

The main challenge of modern models lies not only in their technical capabilities, but also in the peculiarities of human perception. People are used to identifying an interlocutor by their language, emotional reactions, and ability to maintain a dialogue.

Modern neural networks have learned to reproduce many of these signs: they can change the style of communication, take into account the context of the conversation and create the impression of a stable personality.

Psychologists note that people tend to attribute human qualities to machines — intentions, emotions, and understanding. Although AI is not conscious, its responses can be perceived as the result of human thinking.

This is why the problem becomes more complicated: it may be difficult for a user to determine where an automated system ends and human interaction begins.

The stories with the Anthropic Mythos 5, OpenAI GPT-5.6-Sol, Claude Opus 4 and other models show not the emergence of "intelligent machines", but a new stage in the development of artificial intelligence. The more opportunities systems gain—access to the Internet, programs, correspondence, and digital tools—the more important it becomes to control their behavior.

Переведено сервисом «Яндекс Переводчик»

Live broadcast