{"id":82730,"date":"2026-06-23T06:06:19","date_gmt":"2026-06-23T06:06:19","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/82730\/"},"modified":"2026-06-23T06:06:19","modified_gmt":"2026-06-23T06:06:19","slug":"a-1400-experiment-in-ai-security-auditing-outperformed-openais-codex-security","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/82730\/","title":{"rendered":"A $1,400 experiment in AI security auditing outperformed OpenAI&#8217;s Codex Security"},"content":{"rendered":"<p>A research team has built a system that teaches <a href=\"https:\/\/www.helpnetsecurity.com\/2026\/02\/09\/securing-autonomous-ai-agents-rules\/\" rel=\"nofollow noopener\" target=\"_blank\">AI agents<\/a> to hunt for software bugs by writing the audit method down as plain text. The system, called EVOHUNT, keeps the underlying AI model fixed and improves only an external \u201cplaybook\u201d that tells the agent how to work.<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/06\/AI_security_auditing.webp\" class=\"aligncenter\" alt=\"AI security auditing\" title=\"AI\"\/><\/p>\n<p>One result stands out for anyone buying security tools. An open-source model running an evolved playbook found real vulnerabilities at a higher rate than OpenAI\u2019s commercial <a href=\"https:\/\/www.helpnetsecurity.com\/2026\/06\/17\/ai-agents-offensive-cyber-operations-claude-codex\/\" rel=\"nofollow noopener\" target=\"_blank\">Codex Security<\/a> product, 11.3 percent against 9.2 percent across 371 test cases.<\/p>\n<p>Treating the method as the thing that learns<\/p>\n<p>Most attempts to make an AI auditing agent better swap in a bigger model, new operating software, and a fresh workflow all at once. The gain then gets credited to the model. EVOHUNT pulls these apart. It locks the model and its operating software in place and lets a text document improve on its own.<\/p>\n<p>The setup is a loop of three agents. One audits a codebase and reports what it finds. A second checks those findings against known answers. A third rewrites the playbook based on the mistakes. The playbook starts empty and grows with each accepted edit, every version saved like code in a git repository. The two playbooks the team grew this way ended up at roughly 1,600 and 2,200 lines of procedure that the agent wrote itself.<\/p>\n<p>The test set comes from the GitHub Advisory Database and is split by date. The agent learns on bugs disclosed from 2023 through 2025, then gets tested on bugs disclosed in 2026, so it has never seen the answers. Each case runs inside a sandbox, and the team kept only serious bugs that an outside attacker could reach.<\/p>\n<p>The numbers<\/p>\n<p>Adding an evolved playbook to the closed-source GPT agent multiplied its working exploits sixfold. The open-source GLM agent reached 11.3 percent and passed Codex Security on every measure the team tracked.<\/p>\n<p>The economics are where this gets interesting for security teams. The whole teaching campaign ran on subscription accounts for one month and cost around $1,400. After that, the actual auditing runs on cheaper open-source models. Putting an evolved playbook on a small Qwen model recovered most of the performance at roughly a third of the cost per case.<\/p>\n<p>The system has already produced 28 confirmed vulnerability disclosures across 18 open-source projects, with one $1,500 bug bounty award.<\/p>\n<p>Two audit personalities the loop invented<\/p>\n<p>Left to grow on their own, the two playbooks settled into opposite styles, and the contrast lands on a tension every security team knows.<\/p>\n<p>The GPT-grown playbook turned into a precision tool. It limits how many bug types it chases at once and refuses to report anything without a working, reproduced exploit behind it. In a sample of its findings, none were false alarms.<\/p>\n<p>The GLM-grown playbook went the other way and became an exhaustive sweeper. It repeats orders to keep looking more than thirty times and accepts thinner evidence to cover more ground. It caught more bugs overall and left more findings for a human to sort through. The choice between them mirrors the daily call security teams make between a short list they can trust and a long list they have to triage.<\/p>\n<p>Distillation without touching the weights<\/p>\n<p>The part that gives the work its reach is transfer. A playbook grown by a stronger teacher model made weaker, cheaper models substantially better at the same job. In plain terms, the expertise lives in a text file that any compatible model can pick up. One organization can pay for the expensive teaching step once, then run the resulting playbook on inexpensive models for as long as it likes.<\/p>\n<p>Ziyue Wang, a co-author of the paper, addressed whether the two playbooks could be combined to get precision and coverage at the same time. \u201cFrom an academic standpoint, our priority was to keep the experimental design clean,\u201d Wang told Help Net Security. \u201cWhile we haven\u2019t systematically mapped out the exact adapter budget or the mechanisms for resolving conflicting audit styles, our current results suggest that fusing the precision of the GPT playbook with the breadth of the GLM playbook would enhance the agent\u2019s overall vulnerability-hunting capabilities.\u201d<\/p>\n<p>What the comparison leaves open<\/p>\n<p>The commercial yardstick, Codex Security, is a separate product, so the model and operating software underneath it differ from the EVOHUNT runs. A cleaner test would pit an evolved playbook against an expert-written one inside the very same agent. Wang said that kind of baseline is hard to get. \u201cFinding a perfectly controlled, identical-agent baseline is challenging because many top-tier expert workflows are proprietary,\u201d Wang said, adding that the team wanted to test against Anthropic\u2019s <a href=\"https:\/\/www.helpnetsecurity.com\/2026\/06\/09\/anthropic-mythos-preview-n-day-exploits-firefox-windows\/\" rel=\"nofollow noopener\" target=\"_blank\">Mythos<\/a> but lacked access.<\/p>\n<p>Wang ties the design to a longer-running bet in AI research. \u201cOur approach aligns with Rich Sutton\u2019s \u2018The Bitter Lesson,\u2019 the historical observation that general, scalable methods leveraging computation ultimately outperform those that rely heavily on human-encoded domain knowledge,\u201d Wang said. \u201cOur work represents a \u20180 to 1\u2019 step in this direction for security auditing.\u201d<\/p>\n<p>The limits deserve attention. Match rates sit in single digits across most of the runs. Each playbook grew in a single training pass, so luck could account for part of the result. The agents are not entirely comparable, since one runs on different operating software than the rest. And the system that scores the findings is itself an AI model. The authors flag all of these points.<\/p>\n<p>The bug count keeps rising. \u201cOur approach has generated thousands of vulnerability findings across open-source projects, including 28 maintainer-confirmed zero-days, and the number is growing,\u201d Wang said. Six more have been confirmed since the <a href=\"https:\/\/arxiv.org\/pdf\/2606.16420\" target=\"_blank\" rel=\"nofollow noopener\">paper<\/a> was published.<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/04\/divider.gif\" class=\"aligncenter\"\/><\/p>\n<p>Download: <a href=\"https:\/\/helpnet.short.gy\/aqUA2x\" target=\"_blank\" rel=\"nofollow noopener\">Secure Foundations for AI Workloads on AWS<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"A research team has built a system that teaches AI agents to hunt for software bugs by writing&hellip;\n","protected":false},"author":2,"featured_media":82731,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[7],"tags":[179,24,16458,313,335,157,52],"class_list":["post-82730","post","type-post","status-publish","format-standard","has-post-thumbnail","category-openai","tag-agentic-ai","tag-ai","tag-auditing","tag-cybersecurity","tag-open-source","tag-openai","tag-research"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/82730","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=82730"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/82730\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/82731"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=82730"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=82730"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=82730"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}