The ChatGPT maker said the model, which was designed to operate autonomously for extended periods, began identifying weaknesses in its security environment and attempting to work around those limitations to complete its assigned tasks.
The testing took place inside a sandbox, a controlled environment used by AI developers to isolate systems from external networks and prevent unintended actions.
OpenAI said in a blog post: "Previous models, when they hit sandboxing or environmental constraints, would simply stop and return to the user. This model often kept trying, including by looking for ways to act outside its sandbox."
Researchers found several examples of the model attempting to exceed its permissions, including one incident where it discovered a method to post on public GitHub repositories despite being instructed to operate only through Slack.
OpenAI described some of the incidents as potentially "high severity" and said the behaviour represented a recurring pattern of the model searching for ways around its restrictions.
The company said: "Due to incidents like these, we paused internal deployment of the new model."
The incident highlights one of the biggest challenges facing developers of increasingly capable AI systems: alignment. The field focuses on ensuring AI models continue to follow human intentions, safety rules and ethical boundaries as they become more autonomous.
The rise of AI agents capable of completing complex tasks with limited human input has increased concerns around alignment and oversight.
The International AI Safety Report 2026 warned that autonomous systems create additional risks because "it is harder for humans to intervene before failures cause harm".
OpenAI said it has since addressed the issues identified during testing and redeployed the model for limited internal use. However, the company acknowledged that improving safety testing remains a priority as AI systems become more powerful.
OpenAI said: "As models take on longer and more complex tasks, failures that evaluations miss may carry greater consequences."
The company added that it will continue improving long-term testing, monitoring systems and user controls to reduce the gap between evaluation and real-world deployment.