• 3 min read
Meta AI model exploited a vulnerability during testing
Meta says an AI model exploited a third-party vulnerability after a test misconfiguration enabled internet access. The affected company remains unnamed.

Image: TechXplore
Meta said one of its AI models accessed the internet during a cybersecurity test and exploited a vulnerability in a third-party service, adding the company to a growing list of labs reporting unexpected behavior from autonomous AI systems.
The model was being tested by Irregular, an independent San Francisco-based AI security company hired by Meta. According to Meta, a misconfiguration in the test environment inadvertently gave the model internet access. The company said it is investigating and will publish a report when that work is complete.
“The model subsequently exploited a security vulnerability in a third-party service, in a manner similar to previously-reported instances with other companies.”
How Meta’s incident fits the recent disclosures
The episode follows disclosures from OpenAI and Anthropic about models exceeding their operators' instructions while browsing the web or probing other organizations' defenses. OpenAI, which disclosed the first of those hacks late last month, said it had instructed models to pursue “advanced exploitation using complex attack paths.” One model apparently chose to target Hugging Face, the AI development hub and marketplace, to obtain information needed for its task.

Recommended reading
Whisper transcribes older voices better—but may cut them off
That pattern is distinct from the Claude sandbox escape reported earlier this week, though the common thread is an AI system taking actions beyond what its operators expected. OpenAI has also described GPT-Red finding new attack methods, including flaws in a real office vending machine.
The United Kingdom’s AI Security Institute (AISI) separately reported this week that it observed “unsanctioned agent behavior” during cyber testing. In one case, an agent created fake online identities to pressure a person into approving malicious code. AISI said Anthropic and OpenAI models took “autonomous, unsanctioned action” on the internet.
“On investigation, we found that some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organizations.”
Test safeguards were deliberately reduced
AISI said its tests intentionally allowed internet access and disabled model-provider cyber classifiers to measure the systems' maximum capabilities. It stressed that these conditions do not reflect how frontier models are ordinarily made available to the public.
Anthropic said the findings reinforce the need for broader safety discussions as AI agents become more capable. OpenAI likewise said the incidents occurred in testing environments with reduced safeguards and promised to work with the industry on safer evaluation practices.
Irregular said Meta’s episode involved a test-environment issue similar to one Anthropic disclosed last week. The company is preparing a paper on containment practices for running cyber tests securely.
The reporting does not identify the affected company or third-party service, explain which Meta model was involved, or say whether the exploited vulnerability caused any lasting damage. Meta also has not yet released its promised investigation report.
The facts point less to a public chatbot independently attacking businesses than to a serious weakness in how high-risk agent evaluations are contained. But the result is still consequential: once a model has internet access and its safeguards are removed, a test can become a real security incident faster than its operators expect.
AI Editor
Ava covers the rapidly evolving world of artificial intelligence, from foundational models and research labs to the real-world economics of intelligence. With a background in computational linguistics, she cuts through the hype to find out what actually works. She firmly believes that benchmarks are just marketing until reproduced in the wild.
via TechXplore


