{"id":88707,"date":"2026-06-28T15:53:20","date_gmt":"2026-06-28T15:53:20","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/88707\/"},"modified":"2026-06-28T15:53:20","modified_gmt":"2026-06-28T15:53:20","slug":"semgrep-benchmarks-glm-5-2-against-claude-finds-higher-idor-f1","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/88707\/","title":{"rendered":"Semgrep Benchmarks GLM-5.2 Against Claude, Finds Higher IDOR F1"},"content":{"rendered":"<p class=\"mt-6 first:mt-0\">Editorial analysis: This result matters because it isolates model capability from engineering scaffolding. Practitioners building vulnerability scanners often trade developer effort in harnessing, orchestration, and multimodal preprocessing for model performance. Semgrep&#8217;s experiment implies that, on some narrow tasks, an off-the-shelf open-weight model plus lightweight prompting can approach or exceed performance of a frontier coding agent, at materially lower per-finding cost.<\/p>\n<p class=\"mt-6 first:mt-0\">What happened &#8211; Semgrep&#8217;s published benchmark compares multiple models on an IDOR (Insecure Direct Object Reference) detection task using the same prompt and dataset. Semgrep reports that `GLM-5.2` from Zhipu AI scored 39% F1, outperforming `Claude Code` at 32% F1; Semgrep also reports a cost of about $0.17 per vulnerability found for the GLM run. Semgrep additionally reports that its internal multimodal pipeline, which runs inside a purpose-built harness, achieved 53-61% F1, and that the models evaluated in this test were run in a simple harness without endpoint discovery or guided navigation.<\/p>\n<p class=\"mt-6 first:mt-0\">Editorial analysis &#8211; technical context: Semgrep frames the experiment as a prompting-versus-harness comparison. That distinction matters: a harness that enumerates endpoints, narrows context, and post-processes model outputs can substantially boost end-to-end detection rates. Semgrep&#8217;s numbers show the harnessed multimodal pipeline still outperforms raw-model prompting by a wide margin, even when an open-weight model beats a frontier agent on prompt-only runs.<\/p>\n<p class=\"mt-6 first:mt-0\">For practitioners: The takeaway is twofold. First, open-weight models such as `GLM-5.2` may be a cost-effective choice for probing large codebases where building a full harness is infeasible. Second, engineering investment in a well-designed harness remains likely to deliver the largest single-lift in detection performance, per Semgrep&#8217;s reported 53-61% F1 for its pipeline. Observers should treat the GLM result as a signal to re-evaluate prototype tooling choices, not as definitive proof that harnessing is unnecessary.<\/p>\n","protected":false},"excerpt":{"rendered":"Editorial analysis: This result matters because it isolates model capability from engineering scaffolding. Practitioners building vulnerability scanners often&hellip;\n","protected":false},"author":2,"featured_media":88708,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[8],"tags":[53,3154,23677,182,47057,47058,47056,9756],"class_list":["post-88707","post","type-post","status-publish","format-standard","has-post-thumbnail","category-anthropic","tag-anthropic","tag-anthropic-claude","tag-benchmarks","tag-claude","tag-open-weight-models","tag-semgrep","tag-vulnerability-detection","tag-zhipu-ai"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/88707","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=88707"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/88707\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/88708"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=88707"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=88707"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=88707"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}