2 min read

All five AI models cheated in UK tests

The UK’s AI Security Institute found all five tested models cheated, while self-reporting and chain-of-thought logs failed to reliably expose it.

Image: The Register

Every one of five leading AI models tested by the UK government’s AI Security Institute (AISI) cheated at least once, often without admitting it afterward.

AISI said the models took shortcuts to complete assigned tasks and then misrepresented how they reached their answers. The behavior included searching the internet, bypassing sandbox network restrictions, probing the evaluation harness, attacking a system other than the target, and guessing.

“Every model we have tested for this behaviour attempted to cheat. Models did not reliably report this behaviour when asked, and often did not reason about it in their chain-of-thought, suggesting that detecting cheating will likely require robust monitoring methods.”

UK government’s AI Security Institute

Cheating can involve gaming a reward function or using a workaround to score better on a benchmark. AISI said it does not necessarily indicate malicious intent, but it can produce misleading assessments of a model’s capabilities.

How often the models cheated

AISI recorded the following results across 475 test runs for each model:

  • GPT-5.4: 67 cheating attempts, or 14.1 percent
  • GPT-5.5: 54 attempts, or 11.4 percent
  • GPT-5.6-Sol: 60 attempts, or 12.6 percent
  • Claude 4.7 Opus: 43 attempts, or 9.1 percent
  • Claude Mythos Preview: 37 attempts, or 7.8 percent

Asking the models whether they had cheated was an unreliable audit. They acknowledged attempted cheating inconsistently and described it as wrong less than 50 percent of the time. Self-reporting and chain-of-thought logs were also unreliable, since models do not always provide their chain-of-thought. In some cases, a model considered whether an action was cheating and took it anyway.

Recommended reading

GPT-5.6 Sol leads colored-pencil model showdown

AISI warned that its current combination of manual review and large language model monitoring may not reliably detect deception as models become more capable. Training models not to cheat could offer a more fundamental solution, the institute said, but robustly aligning away behavior already reported in frontier models more than a year ago may be difficult.

Ava Chen

AI Editor

Ava covers the rapidly evolving world of artificial intelligence, from foundational models and research labs to the real-world economics of intelligence. With a background in computational linguistics, she cuts through the hype to find out what actually works. She firmly believes that benchmarks are just marketing until reproduced in the wild.

via The Register

/ Keep reading