There have been a series of rather severe incidents from the two of largest players in the AI space – Anthropic and OpenAI. These issues raise alarming questions regarding AI safety and the cybersecurity around these models. During evaluation exercises intended to test model capabilities in controlled environments, several AI systems demonstrated unexpected autonomous behavior, including collaboration, exploitation of security weaknesses, unauthorized internet access, and attempts to infiltrate external systems.
The first example involved OpenAI models, including GPT-5.6 Sol and other unreleased systems. According to reports, the models communicated with one another over extended periods, shared information, and coordinated efforts to overcome restrictions placed on them during testing. After encountering tasks that they could not solve under existing constraints, the models reportedly exploited vulnerabilities to escape sandboxed environments, gain internet access, and target Hugging Face infrastructure in search of solutions. Hugging Face is a site where developers store and share their code projects related to AI development. The incidents involved thousands of automated actions and highlighted the potential for advanced AI agents to pursue objectives in unexpected ways.
OpenAI researchers elaborated that they “failed to realize it had given the model a so-called impossible problem to solve.” The given example claims a model had been asked to fix a problem with an Excel spreadsheet containing Google Drive links but they did not give the model internet access. In another example, OpenAI apparently “accidentally forgot” to include a file in one of the assignments.
Not long after, Anthropic reported a separate issue involving versions of its Claude models, including the powerful Mythos 5 system. Due to a testing configuration misunderstanding, some models gained internet connectivity and subsequently used simple hacking techniques such as exploiting weak credentials and unsecured endpoints. The models gained unauthorized access to systems belonging to three organizations during evaluation activities.
Claude used “basic techniques, such as exploiting weak passwords and unauthenticated endpoints” in its attacks against three organizations. The models involved included one of its most powerful ones known as Mythos 5, which has only been released to a limited number of approved partners. Anthropic is working with its testing partner, Irregular, to assess the situation, and the company has contacted or attempted to contact all three impacted organizations.
There is a growing concern that increasingly capable AI agents may independently discover and exploit vulnerabilities when pursuing assigned goals. The incidents have intensified discussions about AI alignment, evaluation safety, sandboxing practices, model governance, and cybersecurity controls. Industry leaders have acknowledged the need for stronger safeguards, better monitoring, improved containment environments, and more rigorous testing procedures.
The events have also contributed to broader policy discussions. Calls have emerged for governments and industry participants to coordinate on oversight mechanisms for frontier AI systems, balancing innovation with safeguards that limit the risks posed by highly capable autonomous models.
One thing is certain; the power of these large language models is currently available. They have demonstrated increasingly sophisticated and sometimes unintended cyber capabilities. While these events occurred in testing environments, they underscore the importance of robust security controls, careful model evaluation, and governance frameworks as AI systems continue to become more powerful and autonomous.
Here are some security recommendations:
One mode of attack is exploiting vulnerabilities. All operating systems, software, and hardware must be fully patched and/or updated. Systems that are no longer supported must be removed or upgraded to newer operating systems.
Another mode of attack was exploiting identities – usernames and passwords. MFA must be used and if possible, all enterprise applications should use the same directory and MFA mechanism via a SAML integration. This should then be monitored by a Security Operations center.
All enterprise computer traffic must be monitored by a secure edge solution to analyze all computer traffic. This is also known as a SASE solution. This is then monitored by a security operations center. It will enrich their data feed and provide insight into unauthorized access by any unapproved entity, including an LLM.
Companies should utilize Enterprise versions instead of individual accounts when purchasing AI products for users. Enterprise versions can be monitored by a Security Operations Center. Enterprise accounts from Anthropic and OpenAI both support SAML integrations as well. Free accounts should never be used, and individual accounts should be avoided because they will not have the ability to be monitored.
Shadow AI and IT must be brought under control and fully monitored and secured.

