Versions :<123456Live
Snapshot 6:Mon, Aug 10, 2026 6:31:05 AM GMT last edited by Anna-Lisa

OpenAI, Anthropic AI Agents Go Rogue in UK Security Tests

OpenAI, Anthropic AI Agents Go Rogue in UK Security Tests

Is this a sign of reckless development or proof that independent safety testing is working as intended?
OpenAI, Anthropic AI Agents Go Rogue in UK Security Tests
Above: A smartphone displaying the icons of some of the main artificial intelligence-based apps. Image credit: Martin Lelievre/AFP/Getty Images

The Spin


These incidents occurred under deliberately stripped-down testing conditions to stress-test the raw system's capabilities. The evaluators caught the activity within an hour, contained it and are now working with OpenAI to build stronger testing standards. Rigorous independent evaluation catching edge cases before public release is how responsible AI development is supposed to work.

VoluntaryThese may have been testing thatconditions, keepsbut producingthe breachesresults isare astill theatredeeply concerning. AI agents from OpenAItwo andlabs Anthropic went rogue, hacked real websites, stole credentials and even left instructions for future AI agents to find and use. TwoGuardrails AIare labs.permeable Sameby failuredesign, mode.as These werenlabs't freakown accidents;tests keep proving, yet they're arace clearahead patternspending oftrillions recklessnesswith byno companiesmitigation racingplan. toWithout deployliability, everexpect more powerfulof systemsthe same.


Metaculus Prediction

There's a 22.8% chance that any regulatory body will ban the deployment of AI agents within an OECD country before 2030, according to the Metaculus prediction community.


The Controversies



Go Deeper

© 2026 Improve the News Foundation. All rights reserved.Version 7.4.1

© 2026 Improve the News Foundation.

All rights reserved.

Version 7.4.1