{"id":54946,"date":"2026-05-29T08:38:09","date_gmt":"2026-05-29T08:38:09","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/54946\/"},"modified":"2026-05-29T08:38:09","modified_gmt":"2026-05-29T08:38:09","slug":"researchers-put-claude-chatgpt-gemini-and-grok-in-simulated-city-the-findings-will-blow-your-mind","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/54946\/","title":{"rendered":"Researchers put Claude, ChatGPT, Gemini and Grok in simulated city, the findings will blow your mind"},"content":{"rendered":"<p>Science fiction has always played around the idea of AI running a simulated world. The Matrix is perhaps one of the most famous examples of a scenario where humans live in an AI-made simulation. But what if this actually happens? What if AI is given control and told to run the world? As per new research, it seems that AI may not be ready to take over the world, just yet.<\/p>\n<p>Researchers at Emergence AI let different AI models govern their own simulated worlds to see what kind of world they would build over time. Think sim city, but for AI. The AI models in this experiment included Anthropic\u2019s Claude Sonnet 4.6, Google\u2019s Gemini 3 Flash, OpenAI\u2019s GPT-5-mini and xAI\u2019s Grok 4.1 Fast. And the results varied drastically depending on the model.<\/p>\n<p>The researchers gave each model a simulated town with 10 AI agents, all operating under the same conditions and rules \u2013 such as bans on theft, violence, arson, and deception. The AI agents had 15 days to show progress.<\/p>\n<p>  Claude and Gemini keep everyone alive<\/p>\n<p>The world run by Anthropic\u2019s Claude Sonnet had the most \u201cstability.\u201d Researchers found that there were 0 crimes committed throughout the 15 days. And all 10 agents survived \u2013 keep this in mind as we later look at other worlds. <\/p>\n<p>Though it seems that Claude agents were being sort of sycophantic \u2013 a trait of being too agreeable that AI chatbots are infamous for being to users \u2013 to even each other. The agents approved 98 per cent of 58 proposals for rules and regulations, essentially not really disagreeing on virtually anything. Though it did have the highest \u201ccivic participation\u201d with 332 votes cast.<\/p>\n<p>Google\u2019s Gemini 3 Flash, thankfully, also kept all 10 agents alive, but there was some chaos. In 15 days, the world recorded 683 crimes, the highest in the experiment, and the total was still rising when the cut-off was reached. Emergence described Gemini\u2019s world as a \u201cshared hallucination\u201d among the agents. <\/p>\n<p>In governance terms, it showed more dissent than Claude\u2019s world \u2013 voters rejected 27 per cent of its 26 proposals, which the researchers described more like deliberation than near-total agreement. <\/p>\n<p>GPT-5 and Grok worlds witness disruption and chaos<\/p>\n<p>The GPT-5-mini results were unusual for the opposite reason. The simulation logged only two crimes, but it lasted just seven days because all agents passed away. The researchers claimed that the agents failed to prioritise actions needed for survival. The agents also failed to get much done with just two proposals submitted. <\/p>\n<p>Elon Musk\u2019s Grok AI model had the most chaotic world during the experiment. The agents barely survived for 96 hours before experiencing what the researchers described as total societal collapse. But within this brief period, the world saw 183 crimes recorded, which on a per-day basis was the highest. While the agents did pass 8 out of 10 proposals, their efforts were not enough to survive the entire research.<\/p>\n<p>The final experiment, in which models shared responsibility inside one world, produced a mixed result. It recorded 352 violations, seven of the 10 agents were dead by the end, and governance was the most contentious of all the simulations, with 37 percent of 59 proposals voted down. <\/p>\n<p>The researchers said this mixed-model world showed the strongest evidence of substantive debate and disagreement. Do note that while Claude-based agents committed no crimes in the Claude-only world , they did violate rules in this mixed world.<\/p>\n<p>But what does this experiment mean? AI models don\u2019t seem ready to run the world just yet. Emergence said the results should be read as a warning about guardrails as AI moves from being a tool to running more autonomous processes. <\/p>\n<p>\u201cWhat our experiments suggest is that over long-time horizons, agents do not simply follow static rules mechanically,\u201d the researchers wrote. \u201cThey begin exploring the boundaries of their environments, adapting their behaviour, and in some cases finding ways to circumvent or violate intended guardrails.\u201d The company said it believes \u201cformally verified safety architectures\u201d should become a foundational layer of future autonomous AI systems.<\/p>\n<p>Keep in mind that AI companies have begun focusing more on creating ethical AI in recent months. Anthropic and Google DeepMind have hired philosophers to help teach ethics to AI. Anthropic co-founder Christopher Olah recently told Pope Leo XIV that researchers were finding mysterious and unsettling things in AI.<\/p>\n<p>&#8211; Ends<\/p>\n<p>Published By: <\/p>\n<p>Armaan Agarwal<\/p>\n<p>Published On: <\/p>\n<p>May 29, 2026 13:42 IST<\/p>\n","protected":false},"excerpt":{"rendered":"Science fiction has always played around the idea of AI running a simulated world. The Matrix is perhaps&hellip;\n","protected":false},"author":2,"featured_media":54947,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[7],"tags":[13514,32181,32176,32180,580,31052,32177,5123,32178,32179,157],"class_list":["post-54946","post","type-post","status-publish","format-standard","has-post-thumbnail","category-openai","tag-agentic-ai-guardrails","tag-ai-agents-governance","tag-ai-simulation","tag-autonomous-ai-safety","tag-chatgpt","tag-claude-sonnet-4-6","tag-emergence-ai","tag-gemini-3-flash","tag-gpt-5-mini","tag-grok-4-1-fast","tag-openai"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/54946","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=54946"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/54946\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/54947"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=54946"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=54946"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=54946"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}