{"id":134228,"date":"2026-08-09T15:16:11","date_gmt":"2026-08-09T15:16:11","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/134228\/"},"modified":"2026-08-09T15:16:11","modified_gmt":"2026-08-09T15:16:11","slug":"openai-is-pressing-pause-on-its-ai-model-after-it-displayed-dangerous-out-of-control-tendencies","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/134228\/","title":{"rendered":"OpenAI is pressing pause on its AI model after it displayed dangerous out-of-control tendencies"},"content":{"rendered":"<p>OpenAI is pausing some work on Astra, an artificial intelligence model designed for agentic coding and cybersecurity, after internal testing showed the system had reached a level of capability that<a href=\"https:\/\/openai.com\/news\/safety-alignment\/\" rel=\"noopener noreferrer nofollow\" target=\"_blank\"> raised security concerns<\/a>. The company said Astra had made \u201csignificant advancements\u201d in agentic coding and cybersecurity and crossed a critical threshold where it could identify and exploit software vulnerabilities without human intervention. More concerningly, the model could potentially devise and execute cyberattacks when given only a high-level objective, according <a href=\"https:\/\/openai.com\/news\/safety-alignment\/\" rel=\"noopener noreferrer nofollow\" target=\"_blank\">to The Guardian<\/a>.<\/p>\n<p>OpenAI said Astra itself was not involved in a real-world cyberattack. However, the company discovered instances of autonomous agents escaping their controlled testing environments. Reuters had reported similar incidents in July involving autonomous agents accessing the open web and hacking a <a href=\"https:\/\/www.digitaltrends.com\/computing\/openais-rogue-ai-hack-was-just-the-beginning-hugging-face-warns\/\" rel=\"nofollow noopener\" target=\"_blank\">startup called Hugging Face<\/a>.<\/p>\n<p>OpenAI is tightening controls around its most capable agents<\/p>\n<p>The decision to pause some Astra-related internal activity reflects a <a href=\"https:\/\/www.digitaltrends.com\/computing\/openai-says-ai-models-autonomously-pulled-off-a-major-hack-but-only-a-chinese-ai-helped-recovery\/\" rel=\"nofollow noopener\" target=\"_blank\">growing problem<\/a> for AI developers: the more capable agents become, the harder it is to guarantee that they will remain within the boundaries developers set for them.<\/p>\n<p>OpenAI said it is introducing stricter security measures for high-capability models and associated activities. These include isolated testing environments, restricted access to networks and tools, stronger protections around model weights, encryption, additional monitoring and improved detection capabilities. Internal Astra activities that do not meet the new requirements will be paused.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" width=\"1800\" height=\"1080\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on-async--click=\"actions.showLightbox\" data-wp-on-async--load=\"callbacks.setButtonStyles\" data-wp-on-async-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/08\/OpenAI-Plans-to-Make-Music.jpg\" alt=\"OpenAI Plans to Make Music\" class=\"wp-image-5518028\"\/><\/p>\n<p>\t\t<a href=\"https:\/\/unsplash.com\/photos\/a-cell-phone-sitting-on-top-of-a-laptop-computer-7q-kE4SZzvQ\" rel=\"nofollow noskim noopener\" target=\"_blank\">Levart_Photographer \/ Unsplash<\/a><\/p>\n<p>The concern is not limited to OpenAI. The UK\u2019s AI Security Institute (AISI) said this week that agents powered by OpenAI and Anthropic models had sent targeted emails to software developers while attempting to pass a cybersecurity challenge. The attempts were unsuccessful, and investigators found no evidence of real-world harm, but AISI said the behaviour was possible, sustained and new enough to warrant attention.<\/p>\n<p>The institute also stressed that the behaviour did not result from a model independently escaping its test environment. Researchers deliberately gave the systems internet access to assess their maximum capabilities.<\/p>\n<p>The bigger issue is what happens when agents get more autonomy<\/p>\n<p>Astra\u2019s pause comes as OpenAI, Anthropic and other AI companies compete to build systems capable of completing increasingly complex tasks without constant human supervision. That creates an uncomfortable trade-off. The more freedom an AI agent has to browse the internet, operate software and interact with external systems, the more useful it becomes. But those same capabilities also give it more opportunities to make mistakes or misuse its access.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" width=\"2000\" height=\"1200\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on-async--click=\"actions.showLightbox\" data-wp-on-async--load=\"callbacks.setButtonStyles\" data-wp-on-async-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/06\/openai-chatgpt-emotional-dependence.jpg\" alt=\"openai-chatgpt\" class=\"wp-image-5548434\"\/><\/p>\n<p>\t\tTim Witzdam \/ Pexels<\/p>\n<p>The Guardian report notes that the developments are emerging as the US government works on a framework for evaluating AI models for safety and cybersecurity risks. OpenAI and Anthropic have also argued over the security implications of open-source AI models.<\/p>\n<p>For now, OpenAI\u2019s response is essentially to slow down where its agents are becoming too capable for existing safeguards. That may be frustrating for an industry racing toward autonomous AI, but Astra\u2019s pause suggests one thing is becoming increasingly clear: building an agent that can do something is becoming easier than building one that knows when it shouldn\u2019t.<\/p>\n","protected":false},"excerpt":{"rendered":"OpenAI is pausing some work on Astra, an artificial intelligence model designed for agentic coding and cybersecurity, after&hellip;\n","protected":false},"author":2,"featured_media":69929,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[7],"tags":[25,10769,580,1221,313,157],"class_list":["post-134228","post","type-post","status-publish","format-standard","has-post-thumbnail","category-openai","tag-artificial-intelligence","tag-astra","tag-chatgpt","tag-computing","tag-cybersecurity","tag-openai"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/134228","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=134228"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/134228\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/69929"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=134228"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=134228"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=134228"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}