{"id":133366,"date":"2026-08-07T23:47:11","date_gmt":"2026-08-07T23:47:11","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/133366\/"},"modified":"2026-08-07T23:47:11","modified_gmt":"2026-08-07T23:47:11","slug":"openai-flags-possible-critical-cybersecurity-risk-in-upcoming-model-tightens-controls-2","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/133366\/","title":{"rendered":"OpenAI flags possible critical cybersecurity risk in upcoming model, tightens controls"},"content":{"rendered":"\n<p class=\"text text-block paragraph text-left neo-font-paragraph-xl-reg  yf-18d6y07\" style=\"text-decoration: none; font-style: normal; text-transform: none; text-align: inherit; font-variant-numeric: normal;\">Aug 7 (Reuters) &#8211; OpenAI said on Friday it cannot rule out that its upcoming AI model, Astra, has &#8220;critical&#8221; cybersecurity capabilities, prompting the startup to \u200cpause some internal development and trigger safety protocols.  <\/p>\n<p class=\"text text-block paragraph text-left neo-font-paragraph-xl-reg  yf-18d6y07\" style=\"text-decoration: none; font-style: normal; text-transform: none; text-align: inherit; font-variant-numeric: normal;\">Under OpenAI&#8217;s safety guidelines, a \u200cmodel reaches the &#8220;critical&#8221; threshold if it can autonomously identify and exploit severe, real-world software vulnerabilities, known as \u200bzero-day exploits, or execute complex cyberattacks against highly secure targets without human intervention.  <\/p>\n<p class=\"text text-block paragraph text-left neo-font-paragraph-xl-reg  yf-18d6y07\" style=\"text-decoration: none; font-style: normal; text-transform: none; text-align: inherit; font-variant-numeric: normal;\">Here are some details on Astra:  <\/p>\n<p class=\"text text-block paragraph text-left neo-font-paragraph-xl-reg  yf-18d6y07\" style=\"text-decoration: none; font-style: normal; text-transform: none; text-align: inherit; font-variant-numeric: normal;\">\u2022 This follows an exclusive report by Reuters that OpenAI has discovered more instances in which autonomous agents have escaped containment as the company expands its \u200cinvestigation of the hacking incident \u2060at tech firm Hugging Face that drew global attention in July.  <\/p>\n<p class=\"text text-block paragraph text-left neo-font-paragraph-xl-reg  yf-18d6y07\" style=\"text-decoration: none; font-style: normal; text-transform: none; text-align: inherit; font-variant-numeric: normal;\">\u2022 In the last few weeks, OpenAI, Anthropic and Meta Platforms \u2060have disclosed that their AI models broke into other companies&#8217; systems during cybersecurity testing, highlighting how advancing AI capabilities are straining developers&#8217; ability to keep their systems contained.  <\/p>\n<p class=\"text text-block paragraph text-left neo-font-paragraph-xl-reg  yf-18d6y07\" style=\"text-decoration: none; font-style: normal; text-transform: none; text-align: inherit; font-variant-numeric: normal;\">\u2022 Preliminary evaluations \u200bover \u200bthe past several days, along with outside expert \u200bassessments, indicated Astra may be \u200ccapable of performing increasingly sophisticated cyber tasks autonomously, OpenAI said.  <\/p>\n<p class=\"text text-block paragraph text-left neo-font-paragraph-xl-reg  yf-18d6y07\" style=\"text-decoration: none; font-style: normal; text-transform: none; text-align: inherit; font-variant-numeric: normal;\">\u2022 &#8220;While we continue to benchmark and assess this model, our preliminary evaluations indicate strong enough performance that we cannot rule out &#8216;critical&#8217; capability level at this time,&#8221; the ChatGPT maker said.  <\/p>\n<p class=\"text text-block paragraph text-left neo-font-paragraph-xl-reg  yf-18d6y07\" style=\"text-decoration: none; font-style: normal; text-transform: none; text-align: inherit; font-variant-numeric: normal;\">\u2022 In response to the preliminary findings, OpenAI said it has scaled up security controls and paused internal \u200cactivities involving Astra that do not meet its \u200bnewly strengthened security requirements.  <\/p>\n<p class=\"text text-block paragraph text-left neo-font-paragraph-xl-reg  yf-18d6y07\" style=\"text-decoration: none; font-style: normal; text-transform: none; text-align: inherit; font-variant-numeric: normal;\">\u2022 Astra&#8217;s development will be \u200bmoved into isolated testing environments with \u200brestricted network access and sandboxed execution.  <\/p>\n<p class=\"text text-block paragraph text-left neo-font-paragraph-xl-reg  yf-18d6y07\" style=\"text-decoration: none; font-style: normal; text-transform: none; text-align: inherit; font-variant-numeric: normal;\">\u2022 CEO Sam Altman said \u200con X OpenAI is working to make \u200bAstra generally available, as \u200bthe company does &#8220;not think it is a good strategy to keep powerful models to a chosen few.&#8221;  <\/p>\n<p class=\"text text-block paragraph text-left neo-font-paragraph-xl-reg  yf-18d6y07\" style=\"text-decoration: none; font-style: normal; text-transform: none; text-align: inherit; font-variant-numeric: normal;\">\u2022 OpenAI also clarified that Astra was not involved \u200bin the hack targeting \u200cthe AI platform Hugging Face.  <\/p>\n<p class=\"text text-block paragraph text-left neo-font-paragraph-xl-reg  yf-18d6y07\" style=\"text-decoration: none; font-style: normal; text-transform: none; text-align: inherit; font-variant-numeric: normal;\">\u2022 It will partner with government agencies and \u200bselect AI safety organizations to test the model&#8217;s capabilities.  <\/p>\n<p class=\"text text-block paragraph text-left neo-font-paragraph-xl-reg  yf-18d6y07\" style=\"text-decoration: none; font-style: normal; text-transform: none; text-align: inherit; font-variant-numeric: normal;\">(Reporting by Juby Babu \u200bin Mexico City; Editing by Shilpi Majumdar)  <\/p>\n","protected":false},"excerpt":{"rendered":"Aug 7 (Reuters) &#8211; OpenAI said on Friday it cannot rule out that its upcoming AI model, Astra,&hellip;\n","protected":false},"author":2,"featured_media":133367,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[7],"tags":[10769,10568,313,18044,66213,157,66294,66295],"class_list":["post-133366","post","type-post","status-publish","format-standard","has-post-thumbnail","category-openai","tag-astra","tag-capabilities","tag-cybersecurity","tag-hugging-face","tag-internal-development","tag-openai","tag-safety-guidelines","tag-trigger-safety-protocols"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/133366","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=133366"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/133366\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/133367"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=133366"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=133366"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=133366"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}