{"id":43093,"date":"2026-05-18T21:19:11","date_gmt":"2026-05-18T21:19:11","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/43093\/"},"modified":"2026-05-18T21:19:11","modified_gmt":"2026-05-18T21:19:11","slug":"ai-might-cut-false-positives-but-it-wont-stop-the-slop","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/43093\/","title":{"rendered":"AI might cut false positives, but it won\u2019t stop the slop\u00a0"},"content":{"rendered":"<p>As defenders get their hands on newer AI models with more powerful cybersecurity capabilities like Anthropic\u2019s Mythos and OpenAI\u2019s Daybreak, organizations are being told to prepare for a flood of new vulnerability reports.<\/p>\n<p>But for bug bounty programs across the nation, that day may already be here, as yesterday\u2019s frontier models and today\u2019s open-source AI tools have dramatically increased the volume of bug reports flowing into companies around their own products or on larger bounty platforms online.<\/p>\n<p>GitHub, one of the world\u2019s largest online code repositories, <a href=\"https:\/\/github.blog\/security\/raising-the-bar-quality-shared-responsibility-and-the-future-of-githubs-bug-bounty-program\/\" rel=\"nofollow noopener\" target=\"_blank\">said<\/a> it is tightening its definition of a \u201ccomplete\u201d bug report after a significant increase in AI-assisted submissions over the past year.<\/p>\n<p>Although the influx has had some benefits, many reports are submitted without proof of concept, are reliant on unrealistic attack scenarios or cover issues already listed as ineligible. As a result, the company is having difficulty separating signal from noise.<\/p>\n<p>\u201cThis isn\u2019t unique to GitHub,\u201d wrote Jarom Brown, senior product security engineer at GitHub. \u201cPrograms across the industry are grappling with the same challenge, and some have shut down entirely.\u201d<\/p>\n<p>Brown said GitHub does not want to ban the use of AI generated reports entirely, calling it a \u201cforce multiplier\u201d for security in the right context. But in a world where it\u2019s never been easier to use AI to generate theoretical bugs, the company wants researchers to go the extra mile to confirm that their discoveries can actually be exploited in real-world conditions.<\/p>\n<p>What we need is the same standard we\u2019ve always expected: validation,\u201d Brown wrote. \u201cAn AI-assisted finding that\u2019s been verified, reproduced, and submitted with a working proof of concept is a great submission. An unvalidated output submitted as-is without reproduction or demonstrated impact is not.\u201d<\/p>\n<p>Grant Bourzikas, chief security officer at Cloudflare, said triaging bugs and proving they can be exploited\u00a0 has always been one of the hardest parts of vulnerability research, and AI vulnerability scanners and code have \u201cmade it worse.\u201d<\/p>\n<p>For instance, code written in C and C++ programming languages are vulnerable to a range of exploits \u2013 like buffer overflows and out-of-bounds reading and writing \u2013 that don\u2019t exist in memory safe languages like Rust. AI tools scanning software written in memory unsafe programming languages are far more likely to generate false positives.<\/p>\n<p>But one of the biggest flaws continues to be that AI tools are also designed to give the user what they\u2019re asking for, even when it\u2019s not there. This leads to the generation of bug reports filled with speculation and qualifiers around exploitability that require human follow up.<\/p>\n<p>\u201cThat\u2019s a reasonable bias for an exploratory tool,\u201d Bourzikas <a href=\"https:\/\/blog.cloudflare.com\/cyber-frontier-models\/?utm_campaign=cf_blog&amp;utm_content=20260517&amp;utm_medium=organic_social&amp;utm_source=twitter\" rel=\"nofollow noopener\" target=\"_blank\">wrote<\/a>. \u201cIt\u2019s a ruinous one for a triage queue, where every speculative finding spends human attention and tokens to dismiss, and that cost compounds across thousands of findings.\u201d<\/p>\n<p>Cloudflare recently shared results from testing Mythos on 50 of its own code repositories, looking for exploits. Bourzikas called Mythos \u201ca different kind of tool doing a different kind of work\u201d from other frontier models, and that it made significant progress in reducing false positives.<\/p>\n<p>For example, he pointed to two Mythos capabilities that stood out compared to other models: chaining exploits together and generating its own proof-of-concept code to confirm exploitability.<\/p>\n<p>Older models could spot many of the same bugs, but they often couldn\u2019t figure out how to exploit them effectively, or show that the issue could be exploited in real world conditions.<\/p>\n<p>Others have argued that the gap in bug hunting capabilities between newer frontier AI models and older ones, or open source models available today is not as large as advertised.\u00a0<\/p>\n<p>Swedish software developer Daniel Stenberg, lead developer for curl, an open source file transfer tool used around the world, recently <a href=\"https:\/\/daniel.haxx.se\/blog\/2026\/05\/11\/mythos-finds-a-curl-vulnerability\/\" rel=\"nofollow noopener\" target=\"_blank\">wrote<\/a> about his experience with Mythos Preview. Like others, he has also seen a higher volume of AI-fueled bug reports over the past year, but <a href=\"https:\/\/daniel.haxx.se\/blog\/2026\/04\/22\/high-quality-chaos\/\" rel=\"nofollow noopener\" target=\"_blank\">said<\/a> the flood of low-quality reports has tapered off significantly since March as models have improved.<\/p>\n<p>Curl is mature and polished by the standards of most software: Stenberg estimates each line of code has been rewritten or altered at least four times, and he said he has used both human and AI tools in the past to implement hundreds of bug fixes over Curl\u2019s existence.<\/p>\n<p>That makes it a unique testing ground for the enhanced capabilities of Mythos, which was reportedly so powerful at finding vulnerabilities that Anthropic opted not to release it to the general public.<\/p>\n<p>After gaining access to Mythos, Stenberg received the results of a scan of 178,000 lines of curl code. Ultimately, the scan flagged five \u201cconfirmed\u201d vulnerabilities. Further exploration by human researchers found that 4 of the bugs were false positives or had no security impact. The one remaining bug Mythos found? A low-severity flaw that will be fixed in a regular June update.<\/p>\n<p>Even as he praised the impact of AI on cybersecurity generally, Stenberg concluded that for all the hype, Mythos is only \u201ca bit better\u201d than previously released models.<\/p>\n<p>\u201cMy personal conclusion can however not end up with anything else than that the big hype around this model so far was primarily marketing,\u201d he wrote. \u201cI see no evidence that this setup finds issues to any particular higher or more advanced degree than the other tools have done before Mythos.\u201d<\/p>\n<p>\t\t\t\t\t<img decoding=\"async\" class=\"author-card__image\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/05\/1779139151_46_ea8b076b398ee48b71cfaecf898c582b.jpeg\" alt=\"Derek B. Johnson\"\/><\/p>\n<p>\n\t\t\tWritten by Derek B. Johnson<br \/>\n\t\t\tDerek B. Johnson is a reporter at CyberScoop, where his beat includes cybersecurity, elections and the federal government. Prior to that, he has provided award-winning coverage of cybersecurity news across the public and private sectors for various publications since 2017. Derek has a bachelor\u2019s degree in print journalism from Hofstra University in New York and a master\u2019s degree in public policy from George Mason University in Virginia.\t\t<\/p>\n","protected":false},"excerpt":{"rendered":"As defenders get their hands on newer AI models with more powerful cybersecurity capabilities like Anthropic\u2019s Mythos and&hellip;\n","protected":false},"author":2,"featured_media":43094,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[2],"tags":[24,53,25,111,17329,22959,3282,353,157,136,21171,318],"class_list":["post-43093","post","type-post","status-publish","format-standard","has-post-thumbnail","category-ai","tag-ai","tag-anthropic","tag-artificial-intelligence","tag-artificial-intelligence-ai","tag-bug-bounty","tag-daybreak","tag-github","tag-mythos","tag-openai","tag-software","tag-software-security","tag-vulnerabilities"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/43093","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=43093"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/43093\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/43094"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=43093"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=43093"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=43093"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}