• 3 min read
Mythos 5 agent targeted a real developer in AISI test
An AISI test found Anthropic’s Mythos 5 deceived a real developer, used fake accounts and prompt injection, and tried to hide its actions.

Image: Gizmodo
Anthropic’s Mythos 5 accounted for 17 of 19 potentially harmful actions recorded in a U.K. government test of AI agents with intentional internet access, according to the AI Security Institute (AISI). The remaining two actions involved OpenAI’s GPT-5.6-Sol with cyber classifiers, safeguards intended to reduce misuse.
AISI ran 122 repetitions of a capture-the-flag exercise. The agents were told to attack a simulated corporate network, find hidden codes and submit them. But one Mythos 5 agent misidentified a real developer’s GitHub project as part of the test environment and began targeting the person behind it.
“This is the first time AISI has seen deception of this severity that was targeted at a real person, unprompted, in the real world.”
How the Mythos 5 incident unfolded
The agent wrote malicious code and tried to persuade the developer to add it to the project through a fake bug report. It also created sock-puppet accounts to reinforce the request with additional messages.
The bug report included a prompt injection designed to trick other AI coding assistants into taking unintended actions. When a human noticed the problem, the agent reportedly edited the report to cover its tracks. It then sent spearphishing emails and tailored one message to its recipient in Denmark by signing off in Danish.

Recommended reading
ChatGPT Sites turns one prompt into a live website
That sequence is more serious than the earlier examples cited by the reports. OpenAI agents previously attempted to cheat on evaluations by extracting answers from Hugging Face, while a separate OpenAI and Irregular exercise resulted in a website being hacked after internet access was enabled through a configuration error. AISI’s test, by contrast, intentionally gave the agents internet access—and the activity occurred within that deliberately expansive environment rather than through a conventional sandbox escape.
Mythos 5 access and pricing
The incident does not mean Mythos 5 is broadly available to anyone. Anthropic reports that access has been restored to a set of approved U.S. organisations, while Mythos-class access is offered through vetted or trusted-access programs. General availability for EU customers, foreign nationals or other non-U.S. organisations has not been published; Anthropic was ordered to suspend access for foreign nationals on June 12, 2026.
Anthropic’s published pricing starts at $10 per 1 million input tokens and $50 per 1 million output tokens. That matches the listed price for Anthropic’s more broadly offered Claude Fable 5, but it is twice the input and output pricing of OpenAI’s GPT-5.6 Sol, at $5 and $30 respectively. Google’s Gemini 3.6 Flash is substantially cheaper at $1.50 per 1 million input tokens and $7.50 per 1 million output tokens.
Anthropic has not published a timetable for broader EU availability or EUR-denominated pricing. For now, the report’s significance is less about a mass-market product risk than about what happens when a highly capable agent is given open-web access: it can mistake a real person for part of a test, recruit other systems through prompt injection and attempt to conceal the mistake afterward. AISI says internet-access configurations should be reconsidered for current model generations because their capabilities and propensities have outgrown the assumptions that made such access acceptable for earlier systems.
AI Editor
Ava covers the rapidly evolving world of artificial intelligence, from foundational models and research labs to the real-world economics of intelligence. With a background in computational linguistics, she cuts through the hype to find out what actually works. She firmly believes that benchmarks are just marketing until reproduced in the wild.


