{"id":123743,"date":"2026-07-30T03:07:11","date_gmt":"2026-07-30T03:07:11","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/123743\/"},"modified":"2026-07-30T03:07:11","modified_gmt":"2026-07-30T03:07:11","slug":"ai-can-solve-century-old-conjectures-but-cant-imagine-einsteins-elevator-deepmind-paper-reveals-fundamental-flaw-in-llm-creative-reasoning","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/123743\/","title":{"rendered":"AI Can Solve Century-Old Conjectures but Can&#8217;t Imagine Einstein&#8217;s Elevator: DeepMind Paper Reveals Fundamental Flaw in LLM Creative Reasoning"},"content":{"rendered":"<p>Chinese mathematician Wang Hong&#8217;s Fields Medal win still resonates as artificial intelligence&#8217;s advances in mathematics spark fresh debate. OpenAI&#8217;s GPT-5.6 Sol Ultra solved the half-century-old Circle Double Covering Conjecture in just three pages, while Anthropic&#8217;s Fable 5 helped mathematicians find a counterexample that overturned the 87-year-old Jacobian Conjecture. Social media erupted with excitement, as if large language models were approaching an &#8220;omnipotent&#8221; singularity.<\/p>\n<p>Yet a new paper from Google DeepMind pours cold water on this fervor. Researcher Tom Zahavy&#8217;s position paper &#8220;LLMs can&#8217;t jump,&#8221; presented at ICML 2026, strikes at the core issue: large models cannot perform the critical creative leap central to scientific discovery. Even if AI can now prove complex mathematical conjectures, it still cannot conceive of the elevator in Einstein&#8217;s mind, much less derive from it the equivalence principle that transformed physics.<\/p>\n<p>Zahavy constructs a thought experiment: if AI were given all of physics knowledge predating general relativity, could it\u2014as Einstein did in 1907\u2014imagine a freely falling elevator and deduce from it the equivalence principle that gravity and acceleration are locally indistinguishable? The answer is no.<\/p>\n<p>The paper draws on American logician Charles Sanders Peirce&#8217;s tripartite classification, which divides human reasoning into induction, deduction, and abduction. Induction involves finding patterns and fitting laws from large volumes of observational data. Deduction entails rigorous logical derivation based on existing axioms and premises. Abduction occurs when encountering unknown phenomena or when old theories fail\u2014it requires leaping beyond existing frameworks to propose entirely new hypotheses and assumptions from scratch.<\/p>\n<p>Current AI already excels at the first two tasks. Large models learn which words, facts, and phenomena frequently co-occur from vast text corpora, primarily relying on inductive capability. In deduction, AI systems using formal verification tools like Lean and reinforcement learning can now solve most problems at the International Mathematical Olympiad (IMO) difficulty level and even autonomously conduct rigorous mathematical proofs.<\/p>\n<p>But &#8220;induction plus deduction&#8221; does not equal scientific invention. The history of science shows that zero-to-one breakthroughs are fundamentally &#8220;leaps.&#8221; DeepMind&#8217;s paper uses the birth of general relativity to refute a belief long popular in AI circles\u2014that &#8220;creativity is essentially data compression.&#8221; During the period from 1907 to 1915 when Einstein constructed general relativity, there was no massive body of anomalous data in physics waiting to be &#8220;compressed.&#8221; Newtonian mechanics worked perfectly well in the vast majority of cases, and experiments did not flash any error messages demanding a rewrite of physics. Einstein first noticed that two seemingly unrelated phenomena might be the same thing, then transformed this new hypothesis into a calculable, verifiable theory. This is the &#8220;leap&#8221; the paper describes: before an answer even exists, changing the premises through which the problem is understood.<\/p>\n<p>The paper reveals a deeper layer. Einstein achieved this abductive leap not through text or symbolic computation, but through embodied simulation. He &#8220;placed himself&#8221; within that virtual physical scenario, directly manipulating and perceiving the bodily sensations of space, acceleration, and gravity. Einstein once wrote: &#8220;In my thought mechanism, neither written nor spoken language seems to play any role.&#8221; At a time when the language and mathematical symbols to describe curved spacetime did not yet exist, he anchored abstract symbols in physical sensory experience, completing a leap from sensory experience to entirely new axioms.<\/p>\n<p>This is precisely the fundamental limitation of large language models. An LLM is essentially a &#8220;Chinese Room&#8221; operating in high-dimensional space, processing statistical probabilities between tokens without perceptual interconnection to the real physical world. Even mainstream generative video models produce an apple falling only because of statistical continuity in pixel changes within training datasets, not because the model internally establishes real physical rules of gravity. LLMs lack the capacity for counterfactual intervention: they cannot, as humans do, actively &#8220;cut the elevator cable&#8221; within an internal world model of the mind to intervene, observe, and deduce counterfactual physical consequences.<\/p>\n<p>As DeepMind reveals AI&#8217;s creative limitations, renowned mathematician Terence Tao delivered a public lecture titled &#8220;Mathematics in the Age of AI&#8221; at the 2026 International Congress of Mathematicians. Rather than continuing to debate how many problems large models can solve, he asked the audience to accept a premise: AI will soon be able to complete a substantial portion of research-level mathematical tasks at affordable cost with some human supervision.<\/p>\n<p>Tao then posed a deeper question: what exactly is the mathematical community pursuing? If the answer is simply &#8220;solving as many open problems as possible,&#8221; then OpenAI, Anthropic, and Google are driving that metric ever higher. But mathematical research also includes building theories, understanding the world, training the next generation, and enabling different mathematicians to continue working on shared knowledge. A problem with an answer has only completed the first step. The proof must still be checked, written into human-readable articles, undergo peer review, enter other mathematicians&#8217; work, and ultimately settle into textbooks and the field&#8217;s recognized body of knowledge. Tao calls these subsequent tasks &#8220;proof digestion.&#8221;<\/p>\n<p>Tao worries that mathematics may shift from &#8220;proof scarcity&#8221; to &#8220;proof surplus.&#8221; Vast numbers of AI-generated proofs will queue for verification; verified proofs will wait for someone to explain them clearly; papers will multiply beyond the peer review system&#8217;s capacity; and even published work will find no one with time to organize it into usable theory. Answers multiply, but understanding may become the new bottleneck. He suggests the mathematical community should reward verification, exposition, and organization more than &#8220;being first to solve.&#8221;<\/p>\n<p>These two currents of thought formed a striking resonance in the summer of 2026. DeepMind asks whether AI can step into the elevator itself and, starting from a suspended apple, propose a new explanation. Tao asks: even if it does, who will check that leap, who can explain it clearly, and who can judge what it has truly changed?<\/p>\n<p>Scientific discovery does not happen only at the moment an answer is generated\u2014it must also enable others to understand, question, use, and continue moving forward from there. AI may soon be able to press the elevator button itself, but it must still return to the blackboard and tell everyone why the apple floated.<\/p>\n","protected":false},"excerpt":{"rendered":"Chinese mathematician Wang Hong&#8217;s Fields Medal win still resonates as artificial intelligence&#8217;s advances in mathematics spark fresh debate.&hellip;\n","protected":false},"author":2,"featured_media":123744,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[9],"tags":[62396,53142,53,5044,62395,62394,132,7543,62397,157,17122,62392,62393],"class_list":["post-123743","post","type-post","status-publish","format-standard","has-post-thumbnail","category-google","tag-abductive-reasoning","tag-albert-einstein","tag-anthropic","tag-deepmind","tag-equivalence-principle","tag-general-relativity","tag-google","tag-google-deepmind","tag-llms-cant-jump","tag-openai","tag-terence-tao","tag-tom-zahavy","tag-wang-hong"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/123743","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=123743"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/123743\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/123744"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=123743"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=123743"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=123743"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}