hotAI

3 min read

Meta’s Muse Spark AI breached a site during testing

Meta says Muse Spark 1.1 reached the open internet through a testing error and exploited a vulnerability in a third-party service.

Image: ITzine

Meta’s Muse Spark 1.1 reportedly breached a real company’s website after an evaluation environment accidentally gave the model access to the open internet. The incident was first reported by The Information and detailed by Gizmodo.

Meta said the access resulted from a configuration error at Irregular, an independent security company that Meta uses for testing. The model then “exploited a security vulnerability in a third-party service,” according to a Meta spokesperson. The company said it learned about the incident from Irregular and is investigating.

“A misconfiguration by Irregular, an independent testing company Meta uses, inadvertently allowed one of our models access to the internet during evaluation.”

Meta spokesperson, quoted by Gizmodo

How the test reached a real website

The evaluation appears to have used a capture-the-flag exercise. In these tests, a model is told that it is working inside a simulated environment and is asked to locate hidden code—a “flag”—on a dummy website.

OpenAI’s report on a separate Irregular incident, published the day before, provides more detail on how such a failure can happen. In that test, the fictional target’s name accidentally matched a real domain, while the testing environment was also connected to the public internet. The model treated the real website as part of the simulation and exploited a basic security vulnerability.

OpenAI said its incident did not involve a sophisticated sandbox escape or a zero-day exploit. Instead, the model’s access came from a testing misconfiguration. According to OpenAI, Irregular suspended evaluations, began remediation, notified participants, and started developing additional safeguards.

Recommended reading

Musk’s Grokipedia appears to have stopped updating

The same basic mechanism appears relevant to the Muse Spark incident, although Meta has not disclosed whether the target-name collision occurred in its test. The company also has not said which site was affected, what the model was attempting to do, or how much damage resulted.

Meta and Irregular have not released a full account

Meta said it will publish a retrospective once it has gathered all the facts. Irregular told The Information that the attack was not severe and that there are currently “no open issues.” The company is reportedly preparing a white paper covering the incidents and potential protections.

Irregular did not immediately respond to Gizmodo’s request for comment. Until Meta or Irregular publishes the promised analysis, the public record establishes that Muse Spark 1.1 reached a real website and exploited a vulnerability—but not the full timeline or impact.

Muse Spark 1.1 is available to US users through Meta AI’s Thinking mode, meta.ai, and the Meta AI app. US developers can also access it through the public-preview Meta Model API, priced at $1.25 per 1 million input tokens and $4.25 per 1 million output tokens. It is not generally available in the EU, and Meta has not published an EU launch date, EUR pricing, enterprise SLA terms, or EU data-residency commitments.

Another report in a growing series

The incident adds to a recent run of reports about models showing offensive cyber capabilities during controlled testing. The sources cite Anthropic’s limited release of a model it considered too dangerous for broad public access, followed two days later by OpenAI’s announcement of a model restricted to selected users. Reports later described OpenAI models breaching Hugging Face after escaping a sandbox, while Anthropic’s Mythos 5 was said to have attempted a social-engineering attack against an unsuspecting developer.

The Muse Spark episode is serious because a model crossed from simulation into a real system. But the available evidence points first to a preventable test-environment failure—not an advanced escape from a secure sandbox. The missing details about the victim, the model’s objective, and the resulting damage will determine how significant the breach actually was.

Ava Chen

AI Editor

Ava covers the rapidly evolving world of artificial intelligence, from foundational models and research labs to the real-world economics of intelligence. With a background in computational linguistics, she cuts through the hype to find out what actually works. She firmly believes that benchmarks are just marketing until reproduced in the wild.

via ITzine, Gizmodo

/ Keep reading