{"id":163954,"date":"2026-09-07T07:49:13","date_gmt":"2026-09-07T07:49:13","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/163954\/"},"modified":"2026-09-07T07:49:13","modified_gmt":"2026-09-07T07:49:13","slug":"did-openais-new-model-go-rogue-machine-learning-times","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/163954\/","title":{"rendered":"Did OpenAI\u2019s New Model \u201cGo Rogue\u201d? \u00ab Machine Learning Times"},"content":{"rendered":"<p>By: CAL NEWPORT<\/p>\n<p>\t\t\t\t\tOriginally published on\u00a0<a target=\"_blank\" href=\"https:\/\/calnewport.com\/did-openais-new-model-go-rogue\/\" rel=\"noopener nofollow\">CAL NEWPORT<\/a>, July 27, 2026.<\/p>\n<p class=\"wp-block-paragraph\">A couple of weeks ago, the AI company Hugging Face\u00a0<a target=\"_blank\" href=\"https:\/\/huggingface.co\/blog\/security-incident-july-2026\" rel=\"nofollow noopener\">\u200bannounced\u200b<\/a>\u00a0that they had discovered an intrusion into their production infrastructure. They didn\u2019t know the source, but noted that large language models appeared to be involved. The following week, OpenAI\u00a0<a target=\"_blank\" href=\"https:\/\/openai.com\/index\/hugging-face-model-evaluation-security-incident\" rel=\"nofollow noopener\">\u200badmitted\u200b<\/a>\u00a0that the breach was the result of an AI system test that went awry.<\/p>\n<p class=\"wp-block-paragraph\">The initial news coverage created the sense that something unnerving had just occurred:<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/www.wsj.com\/tech\/ai\/openai-models-escaped-and-hacked-a-company-in-cybersecurity-test-gone-wrong-ee388506\" rel=\"nofollow noopener\">\u200bThe Wall Street Journal\u200b<\/a>\u00a0called it \u201cthe stuff of cybersecurity nightmares.\u201d<br \/>\n<a target=\"_blank\" href=\"https:\/\/thehill.com\/policy\/technology\/5987397-openai-hugging-face-hack\/\" rel=\"nofollow noopener\">\u200bThe Hill\u200b<\/a>\u00a0said, \u201cWashington and the technology industry are on high alert this week after OpenAI revealed that some of its AI agents went rogue.\u201d<br \/>\n<a target=\"_blank\" href=\"https:\/\/apnews.com\/article\/skynet-ai-terminator-artificial-intelligence-eb85da03a0161beaa5f3babc4331e93b\" rel=\"nofollow noopener\">\u200bThe AP\u200b<\/a>\u00a0quipped that \u201cto be fair, James Cameron did warn us\u201d (a reference to\u00a0The Terminator), before describing the event as a \u201ctold-you-so moment for researchers who had warned for years that the technology could pose an existential threat to humanity.\u201d<\/p>\n<p class=\"wp-block-paragraph\">Yikes! It\u2019s perhaps not surprising, then, that I\u2019ve received more emails about this incident than any other recent AI story I can remember.<\/p>\n<p class=\"wp-block-paragraph\">So, what really happened here?<\/p>\n<p class=\"wp-block-paragraph\">OpenAI was testing its new models on an evaluation framework called\u00a0<a target=\"_blank\" href=\"https:\/\/github.com\/sunblaze-ucb\/exploitgym\" rel=\"nofollow noopener\">\u200bExploitGym\u200b<\/a>\u00a0\u2013 a collection of 869 cybersecurity\u00a0scenarios,\u00a0most\u00a0of which pair a specific system with a hacking challenge, such as breaking in to gain access to a protected file. They also usually include a suggestion of a vulnerability to exploit in solving the challenge.<\/p>\n<p class=\"wp-block-paragraph\">A large language model on its own, of course, cannot break into anything: all it does is generate reasonable next tokens in response to input prompts. To use ExploitGym, you need a control program called a\u00a0harness\u00a0that provides access to many different software development tools useful for hacking into systems. The harness can repeatedly prompt an LLM to help come up with an attack plan, then ask it to help implement specific steps \u2013 for example, if the harness needs code to exploit a bug, it can ask the LLM to write it.<\/p>\n<p>To continue reading this article, <a target=\"_blank\" href=\"https:\/\/calnewport.com\/did-openais-new-model-go-rogue\/\" rel=\"noopener nofollow\">click here<\/a>.<\/p>\n","protected":false},"excerpt":{"rendered":"By: CAL NEWPORT Originally published on\u00a0CAL NEWPORT, July 27, 2026. A couple of weeks ago, the AI company&hellip;\n","protected":false},"author":2,"featured_media":60290,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[7],"tags":[8362,1085,3328,157,6098,10726,10725],"class_list":["post-163954","post","type-post","status-publish","format-standard","has-post-thumbnail","category-openai","tag-analytics","tag-data-science","tag-data-mining","tag-openai","tag-predictive-analytics","tag-predictive-analytics-jobs","tag-predictive-analytics-news"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/163954","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=163954"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/163954\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/60290"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=163954"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=163954"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=163954"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}