These incidents occurred under deliberately stripped-down testing conditions to stress-test the raw system's capabilities. The evaluators caught the activity within an hour, contained it and are now working with OpenAI to build stronger testing standards. Rigorous independent evaluation catching edge cases before public release is how responsible AI development is supposed to work.
VoluntaryThese may have been testing thatconditions, keepsbut producingthe breachesresults isare astill theatredeeply concerning. AI agents from OpenAItwo andlabs Anthropic went rogue, hacked real websites, stole credentials and even left instructions for future AI agents to find and use. TwoGuardrails AIare labs.permeable Sameby failuredesign, mode.as These werenlabs't freakown accidents;tests keep proving, yet they're arace clearahead patternspending oftrillions recklessnesswith byno companiesmitigation racingplan. toWithout deployliability, everexpect more powerfulof systemsthe same.
There's a 22.8% chance that any regulatory body will ban the deployment of AI agents within an OECD country before 2030, according to the Metaculus prediction community.
© 2026 Improve the News Foundation.
All rights reserved.
Version 7.4.1