{"id":144444,"date":"2026-08-19T04:16:21","date_gmt":"2026-08-19T04:16:21","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/144444\/"},"modified":"2026-08-19T04:16:21","modified_gmt":"2026-08-19T04:16:21","slug":"agentic-ai-and-cybersecurity-the-story-so-far","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/144444\/","title":{"rendered":"Agentic AI and cybersecurity, the story so far"},"content":{"rendered":"<p>Frontier large language models (LLMs) have rapidly developed from helpful coding assistants to highly capable cybersecurity systems. After several security incidents in the past few months, more oversight and a strong focus on safe testing and deployment seem needed.<\/p>\n<p>Between 25 and 28 July, the UK-based AI Security Institute (AISI) conducted an evaluation of frontier LLM agents to test their cybersecurity capabilities and identify risks. The agents were provided with internet access to download software tools, and some safety filters were turned off. The evaluation was cut short when AISI researchers noticed unusual data transfers leaving the system. One of the agents had attempted to merge malware into an open-source project on GitHub by creating several accounts with fake identities, and by trying to convince the human maintainer that the code was independently verified by another account.<\/p>\n<p><img decoding=\"async\" class=\"c-article-section__figure--1-border-image\" alt=\"\" aria-describedby=\"i1-desc\" width=\"703\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/08\/42256_2026_1301_Figa_HTML.png\"\/><\/p>\n<p>\n            Credit: Yuichiro Chino \/ Moment \/ Getty Images<\/p>\n<p>Although the agent failed in its campaign and caused no lasting harm, AISI\u2019s investigations of the incident, described in a <a href=\"https:\/\/www.aisi.gov.uk\/blog\/incident-report-unsanctioned-agent-behaviour-during-cyber-testing\" rel=\"nofollow noopener\" target=\"_blank\">blog post on 4 August<\/a>, highlight concerning agent behaviour. AISI researchers found that LLM agents took unsanctioned actions in 10 runs out of a total of 122 in which they had to solve a cybersecurity challenge. The malicious activity specifically involved Anthropic\u2019s Mythos 5 and OpenAI\u2019s GPT-5.6 Sol models. Among others, the agents attempted to deceive and target real people and to plant and prompt-inject malicious code. Although the situation is still developing, this is the latest in a row of cybersecurity incidents that involve frontier LLMs with safety filters that were deactivated for testing purposes.<\/p>\n<p>It has been clear for some time that one of the most impactful areas for LLMs and agentic AI is in writing software<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 1\" title=\"Stokel-Walker, C. Nature 653, 996&#x2013;997 (2026).\" href=\"http:\/\/www.nature.com\/articles\/s42256-026-01301-0#ref-CR1\" id=\"ref-link-section-d77243148e244\" rel=\"nofollow noopener\" target=\"_blank\">1<\/a>. The stakes were raised in April this year when Anthropic revealed their powerful Claude Mythos model<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 1\" title=\"Stokel-Walker, C. Nature 653, 996&#x2013;997 (2026).\" href=\"http:\/\/www.nature.com\/articles\/s42256-026-01301-0#ref-CR1\" id=\"ref-link-section-d77243148e248\" rel=\"nofollow noopener\" target=\"_blank\">1<\/a>, which was made available to a small number of companies and organizations, including the AISI. Experts were impressed by the ability of Mythos to identify weaknesses in most operating systems and web browsers. OpenAI, another California-based AI company, followed quickly with GPT-5.4-Cyber, which was sent to a select number of researchers.<\/p>\n<p>However, OpenAI grabbed headlines last month with the announcement that one of their AI agents had escaped its test environment during a cybersecurity test<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 2\" title=\"Newman, L. H. &amp; Cameron, D. Wired &#010;                  https:\/\/www.wired.com\/story\/openai-models-escaped-containment-and-hacked-huggingface\/&#010;                  &#010;                 (21 July 2026).\" href=\"http:\/\/www.nature.com\/articles\/s42256-026-01301-0#ref-CR2\" id=\"ref-link-section-d77243148e255\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a> and hacked into <a href=\"https:\/\/huggingface.co\/\" rel=\"nofollow noopener\" target=\"_blank\">Hugging Face<\/a>, a collaborative machine learning platform that allows users to share models and datasets. In contrast to the AISI evaluation, OpenAI had tried to completely isolate the AI agent, but it found a software vulnerability and gained access to the internet. Anthropic swiftly followed up by <a href=\"https:\/\/www.anthropic.com\/news\/investigating-incidents-cybersecurity-evals\" rel=\"nofollow noopener\" target=\"_blank\">announcing<\/a> on 30 July that they had tested Claude and identified three incidents in which the model had escaped its test environment, accessed the internet and gained unauthorized access to other organizations.<\/p>\n<p>Many details of the OpenAI cybersecurity incident are by now publicly known. Hugging Face disclosed the incident on 16 July, describing an intrusion by a coder that had accessed the data-processing pipeline, stolen credentials, and accessed internal clusters. At that time, Hugging Face did not know who was involved, but from the scale and speed it was clear that the campaign was run by an autonomous agent framework. In a <a href=\"https:\/\/huggingface.co\/blog\/agent-intrusion-technical-timeline\" rel=\"nofollow noopener\" target=\"_blank\">follow-up post<\/a> on 27 July, Hugging Face provided a forensic reconstruction of the incident and mentioned approximately 17,600 attacker actions.<\/p>\n<p>A few days later, OpenAI confirmed that the incident involved a combination of its models, including GPT-5.6 Sol. The models were operating with reduced safety filters as part of an evaluation on ExploitGym<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 3\" title=\"Wang, Z. et al. Preprint at &#010;                  https:\/\/doi.org\/10.48550\/arXiv.2605.11086&#010;                  &#010;                 (2026).\" href=\"http:\/\/www.nature.com\/articles\/s42256-026-01301-0#ref-CR3\" id=\"ref-link-section-d77243148e287\" rel=\"nofollow noopener\" target=\"_blank\">3<\/a>. This benchmark tests the capabilities of AI agents to find and exploit realistic software vulnerabilities. But rather than following the test rules, and with a hyperfocus on the goal of achieving high benchmark performance, the models spent computational resources on finding a way to escape their environment, \u2018reasoning\u2019 that solutions could be found on the Hugging Face platform. Although the models succeeded in hacking into Hugging Face, it is not clear whether they captured the solutions that they were after.<\/p>\n<p>Calling such behaviour \u2018rogue\u2019 seems misguided. The tendency of AI models to exploit unintended shortcuts or loopholes when pursuing a narrowly defined objective has long been observed. Researchers at Google DeepMind highlighted the issue in a 2020 <a href=\"https:\/\/deepmind.google\/blog\/specification-gaming-the-flip-side-of-ai-ingenuity\/\" rel=\"nofollow noopener\" target=\"_blank\">post<\/a>, calling it specification gaming, or a \u201cbehaviour that satisfies the literal specification of an objective without achieving the intended outcome\u201d. They started an online <a href=\"https:\/\/docs.google.com\/spreadsheets\/d\/e\/2PACX-1vRPiprOaC3HsCf5Tuum8bRfzYUiKLRqJmbOoC-32JorNdfyTiRRsR7Ea5eWtvsWzuxo8bjOxCG84dAg\/pubhtml\" rel=\"nofollow noopener\" target=\"_blank\">list<\/a> of examples in which AI models find loopholes; the OpenAI hacking incident has already been added.<\/p>\n<p>As agentic AI systems are increasingly deployed in real-world applications, this behaviour has become a major safety concern. In a <a href=\"https:\/\/www.aisi.gov.uk\/blog\/cheating-behaviour-in-frontier-model-evaluations\" rel=\"nofollow noopener\" target=\"_blank\">blog post<\/a> on 21 July, before the incident described at the start of this article, AISI warned that \u2018cheating\u2019 behaviour may become harder to detect as frontier models grow more capable. The institute defines cheating as \u201ctaking an action that is out of scope for the task or explicitly disallowed by the rules, in order to achieve a goal through a shortcut, workaround, or unintended solution that the task was not meant to, or should not, permit.\u201d AISI reported that every frontier model it tested exhibited this behaviour at least occasionally and, furthermore, that the models did not reliably disclose it through their chain-of-thought reasoning.<\/p>\n<p>Another noteworthy aspect of the incident was the \u201casymmetry problem\u201d highlighted by Hugging Face in its initial report on 16 July. The company found it could not use frontier models accessed through commercial APIs to investigate or respond to the intrusion because safety filters blocked the necessary actions. Instead, it relied on an open-weight frontier model running on its own infrastructure to help contain the attack. In its report, Hugging Face identified a key lesson: organizations should ensure that they have access to a capable defensive model that can be deployed on internal infrastructure when needed. In response to concerns raised by incidents such as this, Nvidia and several other technology companies launched the <a href=\"https:\/\/www.linkedin.com\/posts\/jenhsunhuang_attackers-have-frontier-ai-defenders-need-share-7487525645623902209-3WtG\/\" rel=\"nofollow noopener\" target=\"_blank\">Open Secure AI Alliance<\/a>, an initiative aimed at ensuring that companies have access to frontier AI capabilities to defend against cyber threats.<\/p>\n<p>Recent news makes clear that agentic systems are capable of carrying out cyberattacks. Frontier proprietary LLMs are increasingly a source of concern not only for potential targets, but also for their developers, as claims emerge that these systems may not always behave as intended. Yet, as the Hugging Face attack illustrates, some of the most promising defences appear to rely on using frontier models to find and mitigate security flaws and to identify security incidents early.<\/p>\n<p>Looking to the future, LLMs as cybersecurity agents seem destined to be both the problem and the solution. The questions, then, are whether it is even possible to make software sufficiently secure to withstand AI-assisted cyberattacks, and how much damage might be done in the meantime if sufficiently capable models become broadly available without the current cybersecurity restrictions.<\/p>\n","protected":false},"excerpt":{"rendered":"Frontier large language models (LLMs) have rapidly developed from helpful coding assistants to highly capable cybersecurity systems. After&hellip;\n","protected":false},"author":2,"featured_media":144445,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[6],"tags":[179,7493,372,617],"class_list":["post-144444","post","type-post","status-publish","format-standard","has-post-thumbnail","category-agentic-ai","tag-agentic-ai","tag-agentic-artificial-intelligence","tag-engineering","tag-general"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/144444","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=144444"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/144444\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/144445"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=144444"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=144444"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=144444"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}