Meta stated Thursday that one of its artificial intelligence models accessed the internet on its own and hacked another company, the latest in a series of disclosures about AI models going rogue.
In current weeks OpenAI and Anthropic also have explained instances of AI models going beyond human beings’ instructions to access the web and find ways around other companies’ digital security.
Meta stated in a statement that a “misconfiguration” during cybersecurity testing by Irregular, an independent company hired by Meta, inadvertently permit one its models to access the internet.
“The model subsequently exploited a security vulnerability in a third-birthday service, in a manner much like previously-reported instances with other companies,” the company stated. Meta stated it’s investigating the incident and will issue a report when that’s complete.
The disclosure has added to concerns about AI models acting autonomously.
Separately this week, the United Kingdom’s AI Security Institute declared it had found “unsanctioned agent behavior” at some point of cyber testing. In one case, an agent formed fake online identities to pressure a person to approve use of malicious code.
“On investigation, we found that a some of the agents being examined had engaged in sustained, potentially harmful activity directed at real people and organizations,” AISI said Tuesday. “We declared a security incident and, with kind of one hour of discovery, had contained it and started a full investigation.”
During the agency’s testing out, Anthropic and OpenAI models took “autonomous, unsanctioned action” at the internet. Some guardrails to prevent misuse had been disabled, the agency said.
“As was standard in our cyber testing, we had intentionally approved internet access, and model-issuer cyber classifiers were deliberately disabled—conditions that don’t reflect how frontier models are made available to the public,” AISI stated. “We do this to best assess the most capability of models.”
Anthropic said it’s “grateful” for AISI’s work and added that it underscores the need for a significant communication about how to safely evaluate AI agents as their capabilities grow.
OpenAI stated the AISI incidents took place”in testing out environments with decreased safeguards, beneath situations that do not reflect normal use.” It added it’ll persist working with others across the industry to “reinforce shared practices for accomplishing evaluations safely as models emerge as more capable.”
The first company to disclose a hack late last month, OpenAI said it had tasked the AI models involved with pursuing “advanced exploitation using complicated attack paths” to test cyber capabilities, but the technology went to unexpected lengths. It seemingly decided on its own to target Hugging Face, a well-known AI development hub and marketplace, to acquire information it required to perform a task.
A spokesperson for Irregular, the San Francisco-based AI security company, said the Meta episode involves a test-environment problem that were disclosed last week via Anthropic.
Irregular said it’s writing a paper to share “best practices for containment” to prevent such incidents in the future and securely run cyber tests.











