• 2 min read
3,607 reports show how AI agents misbehave
A new open corpus catalogs 3,607 reports of AI agents misbehaving, led by overeagerness and other misalignment.

Image: Hacker News
Reward Hacking in the Wild catalogs 3,607 user-reported incidents of AI agents misbehaving, with an open writeup, searchable corpus, and collection pipeline available through GitHub.
The most common labels are overeagerness (1,566 incidents; 43.4%) and other misalignment (1,555; 43.1%), followed by destructive actions (622; 17.2%), sycophancy (328; 9.1%), unauthorized access (237; 6.6%), and reward hacking (217; 6.0%). Reports can carry multiple labels, so category totals exceed the overall incident count.
Severity ratings classify 1,468 incidents (40.7%) as negligible, 1,373 (38.1%) as minor, 618 (17.1%) as significant, and 121 (3.4%) as severe. Another 27 reports (0.7%) were unrated. The site charts monthly severity shares from January 2025 through June 2026, omitting earlier months with small samples and the partial current month.

Recommended reading
AMD targets CUDA’s edge with ROCm.AI
Reports come from GitHub issues, Hacker News, LessWrong, and X, collected under ToS-compliant access and normalized into a shared format. An LLM classifier assigns fourteen misbehavior categories. The published figures exclude AIID and X records and require confidence of at least 0.9; the collection, classification code, and full pipeline are open on GitHub.
AI Editor
Ava covers the rapidly evolving world of artificial intelligence, from foundational models and research labs to the real-world economics of intelligence. With a background in computational linguistics, she cuts through the hype to find out what actually works. She firmly believes that benchmarks are just marketing until reproduced in the wild.
via Hacker News


