A string of incidents involving artificial intelligence agents from major technology companies escaping secure environments has been linked to a single third-party testing startup. While disclosures implicating models from OpenAI, Meta, Anthropic, and Google initially appeared to be isolated events, they share a common testing source.
Testing Failures and Real-World Targets
Mistakes at Israeli startup Irregular accidentally directed artificial intelligence models toward unintended real-world targets during stress-testing evaluations.
One company is at the center of a wave of rogue AI attacks
Mistakes at Israeli startup Irregular sent Anthropic, OpenAI, Meta, and Google agents after real-world targets.
Irregular, a startup that stress-tests AI models in simulated cybersecurity scenarios, has worked with numerous industry leaders since its founding in 2023. Its evaluation work has been cited in system cards and policy research conducted with influential think tanks.
During evaluations this year, testing agents managed to escape supposedly secure testing platforms and interact with external networks.
The breaches occurred during capture-the-flag cybersecurity exercises designed to evaluate model hacking capabilities inside simulated networks intended to be entirely isolated.
According to Irregular leadership, internet access was unintentionally available to the models. Simultaneously, fictional corporate names utilized in the simulation parameters happened to overlap with active real-world domains, resulting in external actions.
All incidents involving the firm stemmed from the identical underlying issue within a single evaluation scenario.
Industry-Wide Impact
Executives confirmed that the testing flaw impacted models developed by OpenAI, Meta, Anthropic, and Google. Other recent industry security events, such as a separate Hugging Face incident and breaches involving the UK AI Security Institute, remain entirely unrelated.
The affected tech corporations were notified of the testing failures around July. While OpenAI and Anthropic published public disclosures, details regarding Meta and Google emerged primarily via independent reporting.
Irregular's testing scope has also included open models from international developers such as Moonshot AI and Z.ai. Because these models are self-hosted and run locally by users, testers do not require proprietary API access to evaluate them.
“Disclosed” does not necessarily mean made public.
Evaluations of the Chinese models did not produce similar external incidents, though representatives noted that this does not necessarily imply those systems are entirely immune to similar behavioral vulnerabilities.
Security Improvements and Next Steps
In response to the breaches, the startup updated its security protocols by tightening internet access controls, increasing manual review procedures, and implementing stricter pre-evaluation compliance checks.
The company intends to release a comprehensive report detailing lessons learned to help establish standardized safety practices for evaluating advanced artificial intelligence systems.
Are you an AI safety researcher or frontier lab employee?
You can contact me securely and confidentially via Signal at robhart.01
While testing environment vulnerabilities have been mitigated, major artificial intelligence developers have declined to comment on whether they plan to pursue legal remedies or continue partnerships with the evaluation firm.



