{"id":122413,"date":"2026-07-29T06:32:19","date_gmt":"2026-07-29T06:32:19","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/122413\/"},"modified":"2026-07-29T06:32:19","modified_gmt":"2026-07-29T06:32:19","slug":"an-ai-agent-can-pass-every-safety-check-and-still-leak-secrets","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/122413\/","title":{"rendered":"An AI agent can pass every safety check and still leak secrets"},"content":{"rendered":"<p>A pull request lands with a tidy bug report in the description. A bot reads it before any person does, pulls a few shell commands out of it, gets them approved, and posts the output back on the thread. The maintainer reads the whole exchange the next morning.<\/p>\n<p><a href=\"https:\/\/www.linkedin.com\/in\/eladmeged\/\" target=\"_blank\" rel=\"nofollow noopener\">Elad Meged<\/a>, a founding engineer at <a href=\"https:\/\/novee.security\/\" target=\"_blank\" rel=\"nofollow noopener\">Novee Security<\/a>, ran that sequence against three vendors\u2019 own repositories, in the configurations those vendors ship by default. Anthropic\u2019s pipeline handed over secrets. Any organization running one of these agents out of the box carries the same exposure.<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/07\/lock-lines-650.webp\" class=\"aligncenter\" alt=\"AI agent security check\" title=\"Lock\"\/><\/p>\n<p>An agent is a model plus a harness. The model generates intent. The harness turns that intent into shell commands, file reads, API calls, and network requests, and it holds the approval logic, the tool permissions, the path restrictions, and the output handling. When a workflow runs without a person checking each step, the harness becomes the security boundary.<\/p>\n<p>Meged builds AI agents for automated penetration testing at Novee, and he turned the same offensive techniques back onto the agents themselves.<\/p>\n<p>\u201cIn every case, prompt injection was just the delivery mechanism, but the actual vulnerabilities were in how the harness made trust decisions, and how those decisions composed across stages. A command gets approved because it looks safe. The output gets published because that\u2019s the default. Neither decision is wrong alone. Together, they\u2019re an exfiltration chain. If you\u2019re only watching the prompt layer, you never see the handoff where one \u2018safe\u2019 decision feeds into the next,\u201d Meged told Help Net Security.<\/p>\n<p>Every patch drew a new line<\/p>\n<p>Meged reported multiple findings against Claude Code Action, the default workflow configuration for anthropics\/claude-code. Millions of people install the package.<\/p>\n<p>\u201cWe\u2019d report a finding, they\u2019d patch it, and in patching it they\u2019d redraw the line around what counted as safe. The patch itself would tell us where the boundary had moved, which told us exactly where to look next,\u201d he said.<\/p>\n<p>Anthropic awarded bounties across the rounds.<\/p>\n<p>\u201cIt\u2019s important to emphasize that their fixes weren\u2019t inadequate or lazy. Anthropic runs dozens of checks in that pipeline and each patch closed the specific hole, but the problem was that the attacks got harder to detect each round. By the final round, we were recovering secrets through a channel that survived every prior fix: no outbound connection to an attacker, no writes, no logs,\u201d Meged said.<\/p>\n<p>Anthropic paid out every round<\/p>\n<p>\u201cA bounty rewards a finding, scoped to \u2018here\u2019s a specific bug, here\u2019s what it\u2019s worth,\u2019 so a vendor can pay generously round after round and still treat each one as isolated,\u201d Meged said.<\/p>\n<p>\u201cTo be fair, paying out is the right call and they treated us well, so I\u2019m not reading bad faith into it. However, when the same researcher keeps breaking the same architectural seam from different angles, and each fix just moves the boundary to the next surface, the bounty-per-bypass model actually creates a perverse signal: it says \u2018we consider this closed\u2019 when the underlying exposure isn\u2019t. So the framing is important; Google called theirs an \u2018Update to the Trust Model,\u2019 which names the problem structurally.\u201d<\/p>\n<p>Google\u2019s advisory carries a 10.0<\/p>\n<p>Gemini CLI runs on google-gemini\/gemini-cli, a repository with more than 100,000 stars. Meged demonstrated the kill chain there using the security configuration Google\u2019s documentation recommends for CI workflows that process untrusted input. A restriction the operator had configured went unenforced at the point of execution.<\/p>\n<p>The finding came out as part of GHSA-wpqr-6v78-jr5g. The severity sits at the top of the CVSS scale.<\/p>\n<p>OpenAI\u2019s sandbox protects the paths it knows about<\/p>\n<p>Codex CLI ships a default sandbox with every deployment. In multi-stage workflows that share a workspace, state written by one stage gets picked up as trusted context by the next. The protected path list carries a set of assumptions about what needs protecting.<\/p>\n<p>Anthropic built layers of checks into its pipeline, Google built multiple execution modes with environment sanitization, and OpenAI built a sandbox with protected paths. The defenses exist. They fail at the handoffs.<\/p>\n<p>Vendors keep patching surfaces<\/p>\n<p>\u201cA structural fix stops trusting the label and re-validates trust at the point of consumption, not just at the point of decision. Right now, harnesses make a safety call early \u2014 \u2018this command is read-only,\u2019 \u2018this domain is pre-approved\u2019 \u2014 and downstream components inherit that judgment without checking whether it still holds in their context. A read-only command feeding into a public output channel isn\u2019t read-only in effect. A pre-approved domain serving attacker-controlled content isn\u2019t safe in practice,\u201d Meged said.<\/p>\n<p>\u201cA single vendor can absolutely ship meaningful improvements, and some already have after our disclosures. But the pattern repeats across all three vendors we tested, which suggests there\u2019s a shared architectural assumption in how agent harnesses are being built industry-wide. Whether that needs a formal standard or just a shared understanding of the failure mode, the conversation needs to happen across vendors, not just inside one security team.\u201d<\/p>\n<p>Meged will present the code-level analysis and live demonstrations at <a href=\"https:\/\/lp.novee.security\/meet-us-at-blackhat\/\" target=\"_blank\" rel=\"nofollow noopener\">Black Hat USA 2026<\/a>.<\/p>\n<p>What to trace this week<\/p>\n<p>Meged has one audit for teams running these agents in production now.<\/p>\n<p>\u201cTrace every path where the agent\u2019s output, or any state the agent can influence, gets consumed by a later stage with different privileges. Find where your harness says \u2018this is safe\u2019 and then ask: what happens to that output next? Is it published? Is it loaded as configuration? Is it passed to a tool with broader access than the approval assumed?<\/p>\n<p>\u201cMost teams audit what the agent can do. The exposure we keep finding is in what happens after, the handoff between \u2018approved\u2019 and \u2018executed,\u2019 between \u2018read\u2019 and \u2018published,\u2019 between \u2018fetched\u2019 and \u2018trusted.&#8217;\u201d<\/p>\n","protected":false},"excerpt":{"rendered":"A pull request lands with a tidy bug report in the description. A bot reads it before any&hellip;\n","protected":false},"author":2,"featured_media":45078,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[6],"tags":[179,405,53,7537,40923,61544,27026,313,8921,132,30896,8919,720,52],"class_list":["post-122413","post","type-post","status-publish","format-standard","has-post-thumbnail","category-agentic-ai","tag-agentic-ai","tag-ai-agents","tag-anthropic","tag-artificial-intelligence-agents","tag-black-hat","tag-black-hat-usa-2026","tag-conferences","tag-cybersecurity","tag-devsecops","tag-google","tag-novee","tag-penetration-testing","tag-prompt-injection","tag-research"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/122413","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=122413"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/122413\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/45078"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=122413"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=122413"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=122413"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}