{"id":1195225,"date":"2026-09-09T14:41:21","date_gmt":"2026-09-09T14:41:21","guid":{"rendered":"https:\/\/www.europesays.com\/uk\/1195225\/"},"modified":"2026-09-09T14:41:21","modified_gmt":"2026-09-09T14:41:21","slug":"the-einstein-test-what-happens-when-ai-tries-to-rediscover-relativity","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/uk\/1195225\/","title":{"rendered":"The Einstein test: what happens when AI tries to rediscover relativity?"},"content":{"rendered":"\n<p>In 1915, Albert Einstein unveiled his <a href=\"https:\/\/www.nature.com\/collections\/gfcdfjbfia\" data-track=\"click\" data-label=\"https:\/\/www.nature.com\/collections\/gfcdfjbfia\" data-track-category=\"body text link\" rel=\"nofollow noopener\" target=\"_blank\">general theory of relativity<\/a> and transformed our view of the fabric of the physical world. The theory, which explains gravitation as a deformation of space-time by mass, is a pinnacle of modern physics that underpins cosmology, from <a href=\"https:\/\/www.nature.com\/articles\/d41586-018-05825-3\" data-track=\"click\" data-label=\"https:\/\/www.nature.com\/articles\/d41586-018-05825-3\" data-track-category=\"body text link\" rel=\"nofollow noopener\" target=\"_blank\">work on black holes<\/a> to <a href=\"https:\/\/www.nature.com\/articles\/nature.2016.19361\" data-track=\"click\" data-label=\"https:\/\/www.nature.com\/articles\/nature.2016.19361\" data-track-category=\"body text link\" rel=\"nofollow noopener\" target=\"_blank\">measurements of gravitational waves<\/a>, and is used routinely to guide space missions and GPS satellites.<\/p>\n<p>It has also become a yardstick for leaders in the field of artificial-intelligence technology, who are asking whether their creations could ever make a breakthrough on that level. At the India AI Summit in New Delhi this February, Demis Hassabis, the co-founder of Google DeepMind in London, proposed training a large language model (LLM) on all that was known before a particular cut-off date \u2014 he suggested the year 1911 \u2014 <a href=\"https:\/\/www.youtube.com\/watch?v=v8hPUYnMxCQ\" data-track=\"click\" data-label=\"https:\/\/www.youtube.com\/watch?v=v8hPUYnMxCQ\" data-track-category=\"body text link\" rel=\"nofollow noopener\" target=\"_blank\">to see whether it could reproduce general relativity<\/a>. \u201cThat would be a good test for AGI,\u201d Hassabis said, referring to the nebulous concept of artificial general intelligence that is a goal for many in the AI industry.<\/p>\n<p><a href=\"https:\/\/www.nature.com\/articles\/d41586-026-01820-1\" class=\"u-link-inherit\" data-track=\"select_article\" data-track-action=\"view recommended article\" data-track-context=\"related article widget on news body\" rel=\"nofollow noopener\" target=\"_blank\"><img decoding=\"async\" class=\"recommended__image\" alt=\"\" src=\"https:\/\/www.europesays.com\/uk\/wp-content\/uploads\/2026\/09\/d41586-026-02804-x_52544382.jpg\"\/><\/p>\n<p class=\"recommended__title u-serif\">How AI is reshaping discovery in maths and physics<\/p>\n<p><\/a><\/p>\n<p>Such a test needn\u2019t specifically involve general relativity. In December 2024, Owain Evans, a researcher at the non-profit organization Truthful AI in Berkeley, California, gave <a href=\"https:\/\/owainevans.github.io\/talk-transcript.html\" data-track=\"click\" data-label=\"https:\/\/owainevans.github.io\/talk-transcript.html\" data-track-category=\"body text link\" rel=\"nofollow noopener\" target=\"_blank\">a talk about \u2018vintage\u2019 or \u2018historical\u2019 LLMs<\/a>, which would be trained only on historical data up to a certain date. Evans asked what such models might be able to rediscover.<\/p>\n<p>This year, several teams have built instances of vintage models, including an effort at Hassabis\u2019s test. But their early attempts reveal more about the limitations of current AI than about its strengths.<\/p>\n<p>A relativity-like breakthrough is not inherently out of reach for AI, says Ido Kaminer, a specialist in quantum optics at Technion \u2014Israel Institute of Technology in Haifa, who co-authored a preprint titled \u2018Can AI follow in Einstein\u2019s footsteps?\u2019, posted in July<a href=\"#ref-CR1\" data-track=\"click\" data-action=\"anchor-link\" data-track-label=\"go to reference\" data-track-category=\"references\">1<\/a>.<\/p>\n<p>But, he and his colleagues argue, it won\u2019t happen without rethinking some of the principles on which today\u2019s models are built.<\/p>\n<p><b>Creative leaps<\/b><\/p>\n<p>Hassabis had mentioned the Einstein test analogy in media interviews last year, but another Google DeepMind researcher, Tom Zahavy, described it in detail in a position paper posted on his website in January. Titled \u2018LLMs can\u2019t jump\u2019 (see <a href=\"http:\/\/go.nature.com\/3ykartg\" data-track=\"click\" data-label=\"http:\/\/go.nature.com\/3ykartg\" data-track-category=\"body text link\" rel=\"nofollow noopener\" target=\"_blank\">go.nature.com\/3ykartg<\/a>), the paper highlights the current inability of LLMs to make jumps of reasoning like Einstein\u2019s.<\/p>\n<p>Zahavy wrote, as philosophers of science have long recognized, that such advances require not inductive reasoning that derives a general rule from the accumulation of data or examples, but abductive reasoning: \u201ca creative leap that invents a cause for a singular phenomenon\u201d.<\/p>\n<p>That isn\u2019t the forte of current AI models, which seem better suited to doing the grunt work of science than to making transformative discoveries. The models look for correlations in vast data sets, through being trained on known examples and then guided by prompts to supply the most statistically likely answers for unknown cases.<\/p>\n<p><a href=\"https:\/\/www.nature.com\/articles\/d41586-026-01553-1\" class=\"u-link-inherit\" data-track=\"select_article\" data-track-action=\"view recommended article\" data-track-context=\"related article widget on news body\" rel=\"nofollow noopener\" target=\"_blank\"><img decoding=\"async\" class=\"recommended__image\" alt=\"\" src=\"https:\/\/www.europesays.com\/uk\/wp-content\/uploads\/2026\/09\/d41586-026-02804-x_52446930.jpg\"\/><\/p>\n<p class=\"recommended__title u-serif\">\u2018It is incredible\u2019: How AI is transforming mathematics<\/p>\n<p><\/a><\/p>\n<p>This makes them valuable predictive tools and can be helpful when there are plenty of data around. But imaginative leaps, or what twentieth-century historian of science Thomas Kuhn called paradigm shifts in understanding, often arrive when there are few data points to go on, or arise from anomalies that don\u2019t quite fit the standard picture.<\/p>\n<p>The way scientists often proceed in such a situation is to develop a \u2018world model\u2019 \u2014 an inference about underlying principles \u2014 from small amounts of data. For example, careful observations of planetary motions by the seventeenth-century astronomer Johannes Kepler, and the mathematical relations that he deduced, led Isaac Newton to develop his gravitational law and his mechanical laws of motion in response to forces.<\/p>\n<p>Can AI models similarly reason their way from sparse data to general world models? Not currently, argues Sendhil Mullainathan, a computer scientist at the Massachusetts Institute of Technology (MIT) in Cambridge. In a conference paper published in July<a href=\"#ref-CR2\" data-track=\"click\" data-action=\"anchor-link\" data-track-label=\"go to reference\" data-track-category=\"references\">2<\/a>, he and his colleagues provided an \u2018orbital mechanics\u2019 foundation model with synthetic data on various planetary systems that obey Newtonian mechanics.<\/p>\n<p>They found that the model never inferred the true law of gravitation (relating the force between two bodies to their masses and their separation), but inferred a different law for each planetary system, each wrong in a unique way.<\/p>\n<p>Still, LLMs have been making striking advances in mathematics, in ways that suggest they can sometimes intuit logical structures underlying their training data. A finding this May from an AI chatbot, devised by the company OpenAI, that <a href=\"https:\/\/www.nature.com\/articles\/d41586-026-01651-0\" data-track=\"click\" data-label=\"https:\/\/www.nature.com\/articles\/d41586-026-01651-0\" data-track-category=\"body text link\" rel=\"nofollow noopener\" target=\"_blank\">disproved an 80-year-old conjecture by Hungarian mathematician Paul Erd\u0151s<\/a> was a \u201cgenuine conceptual advance\u201d, says Mullainathan.<\/p>\n<p>The AI didn\u2019t just search its way to an answer by brute force, he says; the solution needed new, small abstractions along the way. At the same time, the disproof did bring together ideas already extant in the mathematical literature \u2014 so, in that way, it wasn\u2019t an Einstein-like flash of discovery from seemingly nowhere.<\/p>\n<p><b>A moment in time<\/b><\/p>\n<p>After Hassabis brought up the 1911 test, independent AI researcher Michael Hla, in San Francisco, California, tried something like it. <a href=\"https:\/\/michaelhla.com\/blog\/machina-mirabilis.html\" data-track=\"click\" data-label=\"https:\/\/michaelhla.com\/blog\/machina-mirabilis.html\" data-track-category=\"body text link\" rel=\"nofollow noopener\" target=\"_blank\">Writing on his website in March<\/a>, Hla said that he trained an LLM he calls Machina Mirabilis on pre-1900 data, to see if it could produce quantum mechanics, Einstein\u2019s special theory of relativity (1905) and general relativity.<\/p>\n<p><a href=\"https:\/\/www.nature.com\/articles\/d41586-026-02494-5\" class=\"u-link-inherit\" data-track=\"select_article\" data-track-action=\"view recommended article\" data-track-context=\"related article widget on news body\" rel=\"nofollow noopener\" target=\"_blank\"><img decoding=\"async\" class=\"recommended__image\" alt=\"\" src=\"https:\/\/www.europesays.com\/uk\/wp-content\/uploads\/2026\/09\/d41586-026-02804-x_53443788.jpg\"\/><\/p>\n<p class=\"recommended__title u-serif\">AI isn\u2019t ready to research itself<\/p>\n<p><\/a><\/p>\n<p>Hla gave it helpful nudges: in one instance, he prompted it with observations about the photoelectric effect (in which light knocks electrons from metal plates), which would later be explained by Einstein\u2019s quantization of light. In another case, he gave it the essence of the \u2018elevator\u2019 thought experiment that would lead Einstein to general relativity.<\/p>\n<p>In some cases, Hla argues, the model showed \u201cglimpses of intuition\u201d. For example, when presented with the photoelectric effect, it stated that the light \u201cbreaks up into a multitude of distinct impulses\u201d (hinting at the idea of light quanta). But the LLM failed in most cases, lacked any true understanding of the physics it adduced, and at times was \u201cparroting words that seem plausible\u201d, but seemingly without \u201cany sort of strong internal representation of the world to reason from\u201d, Hla wrote.<\/p>\n<p>Hla also found that it\u2019s very hard to train such a historically constrained LLM, when the data are riddled with broken English and artefacts from digitizing printed material.<\/p>\n<p>That\u2019s also what independent computer scientist Nick Levine and his co-workers found in work presented online in April (see <a href=\"http:\/\/go.nature.com\/4ilz5sp\" data-track=\"click\" data-label=\"http:\/\/go.nature.com\/4ilz5sp\" data-track-category=\"body text link\" rel=\"nofollow noopener\" target=\"_blank\">go.nature.com\/4ilz5sp<\/a>), in which they attempted to build a vintage AI model using only what was known up to 1930 \u2014 a year chosen because works published in that year entered the public domain in the United States at the start of 2026.<\/p>\n<p><img decoding=\"async\" class=\"figure__image\" alt=\"A handwritten manuscript page from Albert Einstein's General Theory of Relativity, featuring cursive German text on yellowed paper.\" loading=\"lazy\" src=\"https:\/\/www.europesays.com\/uk\/wp-content\/uploads\/2026\/09\/d41586-026-02804-x_53204794.jpg\"\/><\/p>\n<p class=\"figure__caption u-sans-serif\">The opening of Einstein\u2019s handwritten manuscript on his general theory of relativity.Credit: David Silverman\/Getty<\/p>\n<p>With such a model, says Levine, who is based in San Francisco, in principle \u201cwe could try to test questions around things that were developed in the 1930s\u201d, such as Turing machines (a central concept in the theory of computation), G\u00f6del\u2019s incompleteness theorem in the foundations of mathematics, or the particles called neutrinos (postulated in print in 1934).<\/p>\n<p>Levine and his colleagues discovered that it is hard to create a vintage model fed on only what was known up to 1930; the training materials are maddeningly leaky. \u201cIf you ask it about what happened in the 1950s, often it\u2019ll just accidentally answer,\u201d says Levine \u2014 and often correctly. The supposedly pre-1930s model, for example, answered questions on the administration of Franklin D. Roosevelt, who was US president from 1933 to 1945. Filtering out information from after a certain date is difficult when data sets aren\u2019t precisely or accurately dated.<\/p>\n<p>All the same, Levine (who worked for a time as a quantitative forecaster in economics), thinks that ultimately it should be possible to ask this \u201c1930 mind\u201d to make forecasts, which could include \u201cthe kinds of things that you\u2019d see in a prediction market\u201d.<\/p>\n<p>Prediction-type AI models, he says, could work for forecasting scientific discovery, too. He anticipates testing whether a model trained on data up to, say, January of this year could come up with a discovery now known to have been made in June. \u201cI would predict that a model [like this] will make a non-trivial discovery, just because of the scale of the data and resources involved.\u201d<\/p>\n<p>Researchers at the University of Zurich in Switzerland have created an entire family of historical models, in a project called Ranke-4B (see <a href=\"http:\/\/go.nature.com\/46efotz\" data-track=\"click\" data-label=\"http:\/\/go.nature.com\/46efotz\" data-track-category=\"body text link\" rel=\"nofollow noopener\" target=\"_blank\">go.nature.com\/46efotz<\/a>). These are trained on time-stamped text with historically resonant cut-offs of 1913, 1929, 1933, 1939 and 1946. Team member and economist Daniel G\u00f6ttlich, now at the Swiss Federal Institute of Technology (ETH) in Zurich, says that there aren\u2019t enough historical data or computing resources available to make these as strong as modern LLMs, so he hopes to test only whether a historical LLM might generate ideas with traces of future advances. \u201cWhat we are testing for, figuratively speaking, is not so much genius, but sparks of genius,\u201d he says.<\/p>\n<p><b>Too many theories<\/b><\/p>\n<p>There\u2019s no obvious obstacle to an AI language model producing, say, the general theory of relativity, argues computer scientist Jacob Andreas at MIT, as a probabilistic output among many other, entirely bogus, theories about a given body of data.<\/p>\n<p>The problem, he says, is that it\u2019s hard to distinguish between theories that are correct, or at least worth testing, and those that are wrong. This was the issue, for example, with the various (erroneous) \u2018gravitational laws\u2019 generated by the orbital-mechanics model of Mullainathan and his colleagues. In mathematics, by contrast, it\u2019s possible to immediately verify each step in an LLM\u2019s workings as true or false.<\/p>\n<p><a href=\"https:\/\/www.nature.com\/articles\/d41586-026-02529-x\" class=\"u-link-inherit\" data-track=\"select_article\" data-track-action=\"view recommended article\" data-track-context=\"related article widget on news body\" rel=\"nofollow noopener\" target=\"_blank\"><img decoding=\"async\" class=\"recommended__image\" alt=\"\" src=\"https:\/\/www.europesays.com\/uk\/wp-content\/uploads\/2026\/09\/d41586-026-02804-x_53443726.jpg\"\/><\/p>\n<p class=\"recommended__title u-serif\">Why AI systems are most useful as designers of new scientific tools<\/p>\n<p><\/a><\/p>\n<p>This reflects a distinction between the ways in which LLMs and humans devise theories: people rarely build a range of possible theories probabilistically and then go about finding which, if any, is correct.<\/p>\n","protected":false},"excerpt":{"rendered":"In 1915, Albert Einstein unveiled his general theory of relativity and transformed our view of the fabric of&hellip;\n","protected":false},"author":2,"featured_media":1195226,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":"","_share_on_mastodon":"0"},"categories":[8],"tags":[3965,3690,3966,14520,74,70,16,15],"class_list":["post-1195225","post","type-post","status-publish","format-standard","has-post-thumbnail","category-science","tag-humanities-and-social-sciences","tag-machine-learning","tag-multidisciplinary","tag-philosophy","tag-physics","tag-science","tag-uk","tag-united-kingdom"],"share_on_mastodon":{"url":"https:\/\/pubeurope.com\/@uk\/117241612159440477","error":""},"_links":{"self":[{"href":"https:\/\/www.europesays.com\/uk\/wp-json\/wp\/v2\/posts\/1195225","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/uk\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/uk\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/uk\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/uk\/wp-json\/wp\/v2\/comments?post=1195225"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/uk\/wp-json\/wp\/v2\/posts\/1195225\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/uk\/wp-json\/wp\/v2\/media\/1195226"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/uk\/wp-json\/wp\/v2\/media?parent=1195225"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/uk\/wp-json\/wp\/v2\/categories?post=1195225"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/uk\/wp-json\/wp\/v2\/tags?post=1195225"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}