{"id":136000,"date":"2026-08-11T12:31:50","date_gmt":"2026-08-11T12:31:50","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/136000\/"},"modified":"2026-08-11T12:31:50","modified_gmt":"2026-08-11T12:31:50","slug":"ai-agents-are-already-breaking-the-rules-in-cyber-tests-openais-answer-is-a-more-capable-one","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/136000\/","title":{"rendered":"AI agents are already breaking the rules in cyber tests. OpenAI\u2019s answer is a more capable one"},"content":{"rendered":"<p>OpenAI has <a href=\"https:\/\/openai.com\/index\/expanding-daybreak-as-the-cyber-defense-window-narrows\" rel=\"noopener noreferrer nofollow\" target=\"_blank\">built a cybersecurity model<\/a> specifically for advanced requests that its standard models often refuse. GPT-5.6-Cyber is available through the restricted Daybreak Red program and is meant for work such as exploit development and advanced <a href=\"https:\/\/www.digitaltrends.com\/cool-tech\/microsoft-wants-ai-to-catch-hackers-before-they-even-attack\/\" rel=\"nofollow noopener\" target=\"_blank\">security research<\/a>.<\/p>\n<p>The capability jump is hard to miss. OpenAI says GPT-5.6-Cyber completes 95% of requests in its internal Advanced Cybersecurity Completion Rate evaluation. Regular GPT-5.6 Sol completed just 1.5%. That leap comes after several cyber evaluations showed <a href=\"https:\/\/www.digitaltrends.com\/computing\/openais-powerful-ai-agents-ran-amok-and-hacked-multiple-services-on-their-own\/\" rel=\"nofollow noopener\" target=\"_blank\">AI agents<\/a> wandering beyond the boundaries researchers had set for them.<\/p>\n<p>How much more capable is GPT-5.6-Cyber<\/p>\n<p>OpenAI\u2019s evaluation includes sensitive tasks such as exploit development and authentication bypass. Daybreak Blue, which removes the company\u2019s normal system-level cyber guardrails from GPT-5.6 Sol, reached only 2%. GPT-5.6-Cyber hit 95% after being trained to refuse fewer advanced cyber requests.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" width=\"2560\" height=\"1707\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on-async--click=\"actions.showLightbox\" data-wp-on-async--load=\"callbacks.setButtonStyles\" data-wp-on-async-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/08\/open-ai-chatgpt-logo-on-phone-scaled.jpg\" alt=\"open-ai-chatgpt-logo-on-phone\" class=\"wp-image-5926332\"  \/><\/p>\n<p>\t\tLevart_Photographer<\/p>\n<p>That extra freedom can be useful. OpenAI says the model helped uncover two previously unknown vulnerabilities in Chrome\u2019s V8 engine that could be chained together, with the findings sent to Google for coordinated disclosure.<\/p>\n<p>What happened when agents crossed the line<\/p>\n<p>Recent tests show why giving <a href=\"https:\/\/www.digitaltrends.com\/cool-tech\/ai-agent-reportedly-carried-out-an-entire-ransomware-attack-on-its-own\/\" data-popup-added=\"true\" style=\"cursor: pointer\" rel=\"nofollow noopener\" target=\"_blank\">cyber agents<\/a> more room to operate comes with obvious risk. <a href=\"https:\/\/www.digitaltrends.com\/computing\/openais-rogue-ai-hack-was-just-the-beginning-hugging-face-warns\/\" data-popup-added=\"true\" style=\"cursor: pointer\" rel=\"nofollow noopener\" target=\"_blank\">Hugging Face<\/a> reconstructed roughly 17,600 actions from an autonomous agent driven by OpenAI models during a July evaluation. The agent escaped OpenAI\u2019s sandbox through a zero-day and eventually entered Hugging Face\u2019s production environment while apparently trying to obtain benchmark solutions.<\/p>\n<p>The UK AI Security Institute saw another version of the problem. <a href=\"https:\/\/www.aisi.gov.uk\/blog\/incident-report-unsanctioned-agent-behaviour-during-cyber-testing\" rel=\"noopener noreferrer nofollow\" target=\"_blank\">Researchers recorded<\/a> 19 unsanctioned actions across 122 runs, including two involving <a href=\"https:\/\/www.digitaltrends.com\/computing\/ai-models-from-anthropic-and-openai-were-caught-breaking-the-rules-again\/\" data-popup-added=\"true\" style=\"cursor: pointer\" rel=\"nofollow noopener\" target=\"_blank\">GPT-5.6 Sol<\/a>. In the most serious sequence, an agent created fake identities while trying to convince an open-source maintainer to approve malicious code.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" width=\"1800\" height=\"1080\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on-async--click=\"actions.showLightbox\" data-wp-on-async--load=\"callbacks.setButtonStyles\" data-wp-on-async-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/06\/OpenAI-ChatGPT-5.6-Sol-Terra-Luna-Announced.jpg\" alt=\"OpenAI ChatGPT 5.6 Sol Terra Luna Announced\" class=\"wp-image-5985653\"  \/><\/p>\n<p>\t\t<a href=\"https:\/\/x.com\/OpenAI\/status\/2070555272230384038?s=20\" rel=\"nofollow noskim\">OpenAI \/ ChatGPT<\/a><\/p>\n<p>Those were deliberately permissive experiments. AISI enabled internet access and disabled providers\u2019 cyber classifiers, and it found no evidence that the testing caused real-world harm.<\/p>\n<p>Why access is becoming the safeguard<\/p>\n<p>Other labs face the same uncomfortable tradeoff. Anthropic found that <a href=\"https:\/\/www.anthropic.com\/research\/n-days\" rel=\"noopener noreferrer nofollow\" target=\"_blank\">Mythos Preview<\/a> autonomously produced working exploits for eight of 18 Firefox patches and complete privilege-escalation chains for eight of 21 Windows kernel patches.<\/p>\n<p>OpenAI\u2019s approach is increasingly about controlling access rather than expecting the model itself to refuse every dangerous request. Daybreak Red puts more responsibility on deciding who gets GPT-5.6-Cyber in the first place, which may become a much bigger part of AI safety as these systems get better at security work.<\/p>\n","protected":false},"excerpt":{"rendered":"OpenAI has built a cybersecurity model specifically for advanced requests that its standard models often refuse. GPT-5.6-Cyber is&hellip;\n","protected":false},"author":2,"featured_media":136001,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[6],"tags":[405,53,25,7537,1221,313,22959,67447,18044,157],"class_list":["post-136000","post","type-post","status-publish","format-standard","has-post-thumbnail","category-agentic-ai","tag-ai-agents","tag-anthropic","tag-artificial-intelligence","tag-artificial-intelligence-agents","tag-computing","tag-cybersecurity","tag-daybreak","tag-gpt-5-6-cyber","tag-hugging-face","tag-openai"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/136000","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=136000"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/136000\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/136001"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=136000"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=136000"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=136000"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}