{"id":124510,"date":"2026-07-30T15:59:15","date_gmt":"2026-07-30T15:59:15","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/124510\/"},"modified":"2026-07-30T15:59:15","modified_gmt":"2026-07-30T15:59:15","slug":"language-models-cant-spark-scientific-revolutions-but-world-models-might","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/124510\/","title":{"rendered":"Language models can&#8217;t spark scientific revolutions, but world models might"},"content":{"rendered":"<p>AI handles two of three types of reasoning<\/p>\n<p>To pinpoint where the gap lies, Zahavy draws on a classic distinction from philosopher Charles Sanders Peirce, who categorized all reasoning by how it connects rules, cases, and results.<\/p>\n<p>Deduction derives guaranteed conclusions from fixed rules, like running a program that produces a provably correct output. Induction spots patterns in data: observe a thousand white swans, and you generalize that all swans are white. Abduction is the creative leap. It invents a cause to explain a surprising phenomenon.<\/p>\n<p>This third form is where Zahavy sees the critical bottleneck, and he draws a line between two levels of it. Ordinary abduction picks the most plausible explanation from a set of known candidates, the way a doctor matches symptoms to a disease. Language models can do this, he concedes. The harder version is what he calls &#8220;manipulative abduction&#8221;: inventing a cause for which no linguistic template exists yet. That, he argues, is the real bottleneck of scientific invention, and machines can&#8217;t do it.<\/p>\n<p>Induction and deduction, the paper argues, are well within reach. Language models already excel at statistical pattern recognition, and they&#8217;re rapidly conquering formal derivation too. Systems like AlphaProof, Gemini, and GPT-5 now achieve gold-level scores on International Mathematical Olympiad problems. Zahavy even concedes that a language model could probably derive general relativity if given Einstein&#8217;s assumptions as a starting point. But formulating those assumptions in the first place, making the manipulative leap to reach them, remains the bottleneck.<\/p>\n<p>Why machines struggle with this leap, Zahavy illustrates using that very theory: AI models typically learn by comparing their predictions to reality and adjusting based on the error, the gap between prediction and outcome. Without a detectable error, there&#8217;s nothing for the system to work with. And that&#8217;s the situation Einstein faced, Zahavy argues.<\/p>\n<p>When Einstein was working, there was no data crisis. Newton&#8217;s physics had been confirmed with extreme precision. The only known anomaly, a tiny shift in Mercury&#8217;s orbit, had been attributed to a hypothetical hidden planet called &#8220;Vulcan.&#8221; An optimization-driven AI would have had no reason to overthrow physics, Zahavy argues. Following the logic of the argument, it would have done what the astronomers of the era did: invented an extra planet to account for the small discrepancy, rather than rethinking space and time. The data confirming Einstein&#8217;s theory, such as Eddington&#8217;s measurement of light deflection, didn&#8217;t arrive until years after the theory was formulated.<\/p>\n<p>A jump requires a body<\/p>\n<p>So where did the manipulative abduction come from that led Einstein to his axioms? Zahavy points to Einstein&#8217;s &#8220;happiest thought&#8221;: the freely falling observer who no longer feels gravity. This insight came from embodied simulation, Einstein mentally playing through a physical sensation rather than grinding through equations. He imagined a physicist inside an accelerating elevator in space and concluded that acceleration and gravity are indistinguishable from the inside.<\/p>\n<p>Zahavy draws a parallel to Archimedes, who didn&#8217;t discover his buoyancy principle through calculation but, as the story goes, through the physical feeling of water rising as he stepped into a bathtub. In both cases, a foundational principle emerged that didn&#8217;t yet exist in the language of the time.<\/p>\n<p>Language models lack exactly this sensory grounding. Zahavy compares them to philosopher John Searle&#8217;s &#8220;Chinese Room,&#8221; a famous thought experiment where a person shuffles Chinese characters according to a rulebook without understanding a single word. Language models shuffle the symbols of physics in much the same way, without access to the physical experience that gives those symbols meaning.<\/p>\n<p>Sakana&#8217;s AI Scientist and Deepmind&#8217;s AlphaEvolve automate scientific workflows impressively. But the AI Scientist only recombines existing concepts, while AlphaEvolve optimizes brilliantly yet needs a clear error signal it can shrink step by step. Einstein never had that signal. Neither system, Zahavy argues, can make the leap into an entirely new framework of thought.<\/p>\n<p>World models as a path to abduction<\/p>\n<p>As a possible way forward, Zahavy points to physically consistent world models. He draws a line here: video generators like <a href=\"https:\/\/the-decoder.com\/deepmind-says-video-models-for-visual-tasks-could-become-what-llms-are-for-text-tasks\/\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">Veo<\/a>\u00a0simply predict the most likely next frame. A falling apple falls not because the model understands gravity, but because falling is the most common continuation in the training data. That&#8217;s still just pattern matching.<\/p>\n<p>Action-controllable world models like\u00a0<a href=\"https:\/\/the-decoder.com\/google-deepmind-opens-project-genie-to-us-subscribers-for-real-time-ai-world-generation\/\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">Genie<\/a>, on the other hand, let an agent actively intervene in a simulation and run counterfactual experiments, like mentally cutting an elevator cable. A &#8220;synthetic lab&#8221; like this could provide the feedback loop needed to invent new axioms where no linguistic template exists yet.<\/p>\n","protected":false},"excerpt":{"rendered":"AI handles two of three types of reasoning To pinpoint where the gap lies, Zahavy draws on a&hellip;\n","protected":false},"author":2,"featured_media":124511,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[9],"tags":[6744,18497,5044,132,7543,10857],"class_list":["post-124510","post","type-post","status-publish","format-standard","has-post-thumbnail","category-google","tag-agi","tag-ai-and-science","tag-deepmind","tag-google","tag-google-deepmind","tag-world-models"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/124510","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=124510"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/124510\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/124511"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=124510"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=124510"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=124510"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}