{"id":84102,"date":"2026-06-24T06:19:09","date_gmt":"2026-06-24T06:19:09","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/84102\/"},"modified":"2026-06-24T06:19:09","modified_gmt":"2026-06-24T06:19:09","slug":"praxen-open-source-ai-agent-behavior-verification","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/84102\/","title":{"rendered":"Praxen: Open-source AI agent behavior verification"},"content":{"rendered":"<p>Praxen is an open-source tool with a simple job: it checks whether an <a href=\"https:\/\/www.helpnetsecurity.com\/2026\/06\/03\/agent-threat-rules-ai-detection\/\" rel=\"nofollow noopener\" target=\"_blank\">AI agent<\/a> does what it claims to do. The tool takes an agent\u2019s declared policy, looks at how the agent operates, and points out every spot where the two drift apart.<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/06\/praxen-AI_agent_behavior_verification.webp\" class=\"aligncenter\" alt=\"Praxen agent behavior verification\" title=\"Open-source AI agent behavior verifier\"\/><\/p>\n<p>It is the reference implementation of Agent Behavior Verification, a control model that hands each agent an authorized role and then confirms the controls hold that agent to it. The idea borrows from how companies manage their own employees. Every person gets a defined set of permissions, and the same logic now applies to software agents, where each one carries a scope of activity it is allowed to perform.<\/p>\n<p>How the verification works<\/p>\n<p>A team writes a Worker Remit, a markdown policy document that declares what the agent may do, including its mission, authorized tools, approved channels, counterparties, and forbidden actions. Praxen then reads evidence such as source code, deployment state, behavioral logs, and governance documents, and reports the gap between declared intent and observed behavior. Findings arrive as a self-contained HTML report, a machine-readable JSON file, and a plain-text summary written to a local reports folder. The tool keeps all data local. Teams install it as a <a href=\"https:\/\/www.helpnetsecurity.com\/2026\/06\/17\/ai-agents-offensive-cyber-operations-claude-codex\/\" rel=\"nofollow noopener\" target=\"_blank\">Claude Code<\/a> plugin.<\/p>\n<p>Each analysis runs a set of named checks. These cover policy-implementation divergence, credential exposure, configuration gaps, capability drift, supply-chain risk, half-wired controls, empty stub files in security-relevant paths, secondary prompt discovery, and compound signal reasoning that chains individual findings into a higher-severity attack path.<\/p>\n<p>Every finding carries tags from the OWASP Top 10 for LLM Applications 2025, the OWASP Top 10 for Agentic AI Applications 2026, the OWASP Secure MCP Server Development Guide 2026, and the RAISE Framework, which assigns a maturity score across six categories. Praxen runs before deployment and on each release. It requires a coding agent, tested against Claude Code, and Python 3.9 or later.<\/p>\n<p>One policy across the agent lifecycle<\/p>\n<p>Runtime monitoring sits in a separate layer called Agent Behavior Analytics. <a href=\"https:\/\/www.linkedin.com\/in\/wilsonsd\/\" target=\"_blank\" rel=\"nofollow noopener\">Steve Wilson<\/a>, Chief AI Officer at Exabeam, told Help Net Security that the company wants the Worker Remit to serve both stages. \u201cOur aim is a single policy,\u201d he said. The remit gives \u201ca structured, human-readable definition of an agent\u2019s intended role, permissions, responsibilities, constraints, and approval requirements.\u201d<\/p>\n<p>Wilson connected the runtime layer to the same definition. \u201cExabeam ABA is designed to analyze the behavior of deployed agents over time and identify activity that deviates from expectations, policy, or established baselines,\u201d he said, adding that the remit \u201cprovides a natural foundation for that analysis because it captures the organization\u2019s explicit expectations for the agent.\u201d Verification answers whether a team built the agent it intended, and analytics answers whether the agent behaves as intended in production. The two capabilities stand separate at present. \u201cOver time, we expect them to become increasingly connected as part of a broader Behavior Intelligence strategy for AI agents,\u201d Wilson said.<\/p>\n<p>Consistency across repeated runs<\/p>\n<p>A <a href=\"https:\/\/www.helpnetsecurity.com\/2026\/03\/13\/claude-code-openai-codex-google-gemini-ai-coding-agent-security\/\" rel=\"nofollow noopener\" target=\"_blank\">coding agent<\/a> performs the analysis, so two runs against the same evidence can produce a different set of findings. Wilson said the major results hold steady. \u201cThe major findings and overall security themes are highly stable,\u201d he said, with smaller movements in severity counts or maturity scoring at the margins. Every finding traces back to source material. Praxen \u201ccites the files, configurations, and artifacts that support its conclusions, allowing a reviewer or auditor to independently verify the claim.\u201d<\/p>\n<p>Exabeam measures consistency with a frozen regression suite of representative agent implementations that validates major findings, themes, and maturity assessments across releases. For governance, compliance, or benchmarking, Wilson recommended that teams \u201crun the analysis multiple times, report the median result and range, and union the material findings across runs.\u201d A single run gives a useful read on an agent\u2019s security posture, and repeated runs add statistical confidence.<\/p>\n<p>Handling evidence that exceeds the context window<\/p>\n<p>Large evidence sets can exceed a model\u2019s context window. Praxen begins with a discovery pass across source code, configuration files, dependency manifests, tool and MCP definitions, memory artifacts, and logs, and prioritizes the material most relevant to agent behavior and security controls. Large logs are sampled to widen coverage.<\/p>\n<p>Wilson pointed to a risk in long-running analysis, where earlier observations can be summarized away as a session grows. Praxen writes findings incrementally and checkpoints the analysis state into a structured manifest before the report is generated. \u201cIf the underlying AI session exceeds its context window, the report can be reconstructed from that checkpoint,\u201d he said. Coverage is recorded directly, so findings drawn from sampled evidence carry a marker and missing evidence registers as a signal of its own. \u201cContext-window limits are a real constraint for every AI-powered analysis platform,\u201d Wilson said. \u201cThe goal is to make them visible, measurable, and recoverable so users can trust the results they receive.\u201d<\/p>\n<p>Praxen is available for free on <a href=\"https:\/\/github.com\/open-agent-ai-security\/praxen\" target=\"_blank\" rel=\"nofollow noopener\">GitHub<\/a>.<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/04\/divider.gif\" class=\"aligncenter\"\/><\/p>\n<p>Must read:<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/04\/devider.webp\"\/><\/p>\n<p>Subscribe to the Help Net Security ad-free monthly newsletter to stay informed on the essential open-source cybersecurity tools. <a href=\"https:\/\/www.helpnetsecurity.com\/newsletter\/\" rel=\"nofollow noopener\" target=\"_blank\">Subscribe here!<\/a><\/p>\n<p><img decoding=\"async\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/04\/devider.webp\"\/><\/p>\n","protected":false},"excerpt":{"rendered":"Praxen is an open-source tool with a simple job: it checks whether an AI agent does what it&hellip;\n","protected":false},"author":2,"featured_media":84103,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[6],"tags":[179,405,7537,12739,3282,335,136],"class_list":["post-84102","post","type-post","status-publish","format-standard","has-post-thumbnail","category-agentic-ai","tag-agentic-ai","tag-ai-agents","tag-artificial-intelligence-agents","tag-exabeam","tag-github","tag-open-source","tag-software"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/84102","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=84102"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/84102\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/84103"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=84102"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=84102"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=84102"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}