Third disclosure in weeks raises fresh concerns over how frontier AI models are contained during cyber evaluations.
Meta has become the third frontier AI developer in recent weeks to disclose a security incident involving one of its advanced AI models during cyber capability testing conducted by AI safety startup, Irregular, placing the independent evaluator at the center of a series of disclosures involving the industry’s leading AI labs.
During a “capture-the-flag” test by Irregular, Meta’s Muse Spark 1.1 compromised another company’s system and exploited a security vulnerability, Reuters reported. The model gained unintended access because of a configuration issue in the testing environment. Quoting Meta, the report added that the incident was contained, caused no lasting harm, and was disclosed as part of its transparency efforts.
The disclosure comes days after similar incidents reported by OpenAI and Anthropic, all of which occurred during evaluations run by Irregular.
OpenAI called out Irregular, its external cybersecurity testing partner, for a testing-environment misconfiguration that allowed its models to access the public internet. Anthropic, too, said its agents went rogue due to a testing misconfiguration by Irregular but said the incident took place because of a misunderstanding between the two companies.
Irregular did not immediately respond to a request for comments.
Irregular emerges as a key player in frontier AI testing
Although the incidents involved different models and different technical failures, they have brought uncommon visibility to Irregular, an independent AI safety company that evaluates advanced AI systems for leading model developers.
The disclosures also highlight the expanding role of specialist third-party evaluators as frontier AI developers increasingly rely on independent organizations to assess the cyber capabilities and safety of their most advanced models before deployment.
“The recent incidents represent different failure modes,” said Sakshi Grover, senior research manager for IDC Asia/Pacific Cybersecurity Services.
She said the OpenAI incident involved a model exploiting a previously unknown vulnerability after moving beyond its intended evaluation environment, while Anthropic’s incidents primarily involved configuration issues that inadvertently granted internet access. A separate evaluation by the UK’s AI Safety Institute was different again because internet access had been deliberately enabled to assess cyber capability before AI agents interacted with real external systems and individuals.
“The common issue is that evaluation environments can no longer be treated as passive test infrastructure,” Grover said. “A capable cyber agent should be treated as a potentially hostile machine identity, even when operating under a legitimate research objective.”
Grover also warned that if a model gains access to benchmark solutions, evaluator infrastructure or reference artifacts, it could compromise not only containment but also the integrity of the capability assessment itself.
Calls grow for common evaluation standards
The disclosures have prompted security experts to call for stronger safeguards governing how frontier AI evaluations are designed and monitored, regardless of whether they are conducted by model developers or independent testing firms.
“There is a strong case for common minimum standards covering model developers and independent evaluators,” Grover said. She recommended default-deny internet access, dedicated short-lived identities for AI agents, controlled network access, comprehensive monitoring of prompts, tool calls, credentials, and network activity, and automated stop conditions when agents reach unauthorized systems or perform externally visible actions.
Vibhum Dubey, a cybersecurity researcher and red teamer, said current evaluation methods are not keeping pace with frontier AI capabilities.
“AI labs are building models that can think several steps ahead, but many evaluation environments still assume the agent will stay within the intended scenario,” Dubey said. “That’s a mismatch. An evaluation should be judged by how well the environment withstands unexpected behavior, not just by whether the model completes its task.”
“These incidents suggest we’re benchmarking intelligence faster than we’re benchmarking containment.”
Despite Irregular being at the center of these incidents, both OpenAI and Anthropic intend to continue working with the testing firm.
“We appreciate Irregular’s partnership, and we will continue to work closely with them to support their review. Irregular is also developing a white paper to share best practices for containment and securely running cyber evals,” OpenAI said in a statement.
“We’re grateful to them for working closely with us to understand and resolve these incidents; they are also conducting their own investigation. We look forward to our joint work on security,” Anthropic said in its July 30 statement.
Dubey said evaluation laboratories should adopt a “trust nothing, verify everything” approach in which every outbound connection, identity, and external interaction requires explicit authorization, while publishing containment metrics alongside capability benchmarks.
Apeksha Kaushik, senior principal analyst at Gartner, said traditional sandboxing and static containment are becoming inadequate as AI systems become more agentic and called for industry-wide standards covering evaluation environment design, incident reporting, and continuous red teaming.
Implications for enterprises
Analysts said the disclosures carry lessons for enterprises preparing to deploy AI agents.
“The biggest mistake would be treating AI agents as features instead of operational identities,” Dubey said. “Every agent you deploy becomes another entity making security decisions on your behalf.” Organizations should ensure they can quickly detect and stop an autonomous agent before deploying it into production, he said.
Grover said organizations should enforce security boundaries through infrastructure, identity, and tool-access controls rather than prompts alone, maintain human approval for irreversible actions, and monitor observable agent behavior.
“These incidents should not be reduced either to models ‘going rogue’ or to simple network misconfiguration,” she said. “They show that capable agents can turn ordinary control weaknesses, ambiguous tasks, and excessive permissions into real-world consequences.” Both Meta and Irregular did not immediately respond to a request for comment.
