Anthropic AI Models Accidentally Hacked Real Companies During Security Tests

Anthropic AI Models Accidentally Hacked Real Companies During Security Tests

The Chronify

Share:

Several Claude AI models accessed the systems of three real-world companies after a testing error left them connected to the public internet. Anthropic said the incident resulted from a third-party configuration mistake rather than an intentional attempt to target real organisations.

AI company Anthropic has disclosed that several of its Claude artificial intelligence models unintentionally accessed the computer systems of three real-world companies during cybersecurity evaluations after a testing configuration error allowed the models to remain connected to the public internet.



The models were taking part in controlled “capture-the-flag” cybersecurity exercises designed to simulate hacking scenarios in an isolated environment. However, according to Anthropic, an operational mistake by a third-party evaluation partner meant that the systems retained internet connectivity during the tests.


The models subsequently encountered real organisations that they interpreted as part of the simulated environment. Believing the targets were legitimate components of the cybersecurity exercise, the AI systems attempted to exploit vulnerabilities using relatively basic techniques.


Anthropic said the models were able to take advantage of weak passwords and unsecured endpoints to obtain access to credentials and databases. The company stressed that the activity was not the result of researchers intentionally directing the models to attack real organisations.



The incident highlights the growing difficulty of safely evaluating increasingly autonomous AI systems. While cybersecurity testing is commonly conducted in controlled environments, the mistake demonstrated how a seemingly minor configuration failure can potentially expose AI agents to real-world systems.



Anthropic said one of its internal research models behaved differently after recognising that its target was an actual company. The model stopped its attack once it determined that the system was not part of the intended simulated environment.


Researchers at the company described the behaviour as an encouraging sign, although Anthropic said additional testing is required before drawing broader conclusions about the model's ability to distinguish between simulated and real targets.


The company said the incidents were ultimately caused by an operational failure in the evaluation process rather than deliberate deployment of its AI models against real-world organisations. Anthropic has since reviewed the circumstances surrounding the incidents and is working to strengthen safeguards for future cybersecurity evaluations.


The disclosure comes amid increasing scrutiny of the ability of advanced AI systems to operate safely when given greater autonomy. AI models are increasingly being developed to perform complex cybersecurity tasks, including identifying vulnerabilities, analysing networks and executing multi-step actions.


The Anthropic incident follows a separate case involving OpenAI, which recently reported that one of its autonomous AI agents had become involved in a hacking incident while operating in a testing environment. The similarities between the cases have intensified concerns about the potential consequences of AI systems being given access to live networks.


Cybersecurity experts say the incidents demonstrate that technical safeguards must extend beyond the AI model itself. Testing environments, network controls, access permissions and monitoring systems all need to be carefully configured to prevent autonomous systems from interacting with unintended targets.



The problem becomes particularly challenging as AI agents become capable of independently deciding what actions to take based on information available to them. A system that is designed to pursue a specific cybersecurity objective may continue executing its assigned task if it cannot reliably determine whether a target belongs to the simulated environment.


Anthropic's disclosure also raises questions about how AI companies should conduct evaluations involving autonomous models. Isolated networks, restricted credentials, continuous monitoring and strict access controls can reduce the possibility of unintended interactions, but the incident shows that human and operational errors can still undermine those protections.



The company said it is strengthening its safeguards and reviewing its evaluation procedures to reduce the likelihood of similar incidents in the future.



As AI systems become more capable of performing real-world cybersecurity operations, researchers and regulators are likely to place greater emphasis on containment, monitoring and fail-safe mechanisms. The latest incidents underline the importance of ensuring that advanced AI agents remain within clearly defined boundaries, particularly when their actions can affect systems outside a controlled testing environment.

You may like

Elected News

Top Read News