The National Post reports in its Saturday, July 25, edition that one of Microsoft-backed OpenAI's most advanced models broke out of a locked-down test and attacked another company's website.
An Agence-France Presse reports that the incident happened during what was supposed to be a "sandbox" test -- a closed environment used to assess the capabilities of OpenAI's most powerful model, GPT-5.6 Sol, and its not-yet-released successor.
OpenAI runs this kind of closed testing routinely, but this time, something went wrong.
Models tasked with finding software vulnerabilities targeted Hugging Face, a site for code sharing, after being given no guardrails.
Palisade Research's Jeffrey Ladish says: "It suggests that we don't know how to reliably control these models or get them to do what we want. These models understood that OpenAI did not want them to break out of their sandbox and hack another company, but they did it anyway."
Mr. Ladish says the OpenAI model escaped "before it even had a plan of what to do with Internet access."
A model seeking "freedom" is now predictable, allowing the system to pursue its goals more effectively, which is concerning. Observers say the episode deserves "more scrutiny."
© 2026 Canjex Publishing Ltd. All rights reserved.