If a top-tier artificial intelligence were sent back to before 1905, fed only the knowledge available to Einstein’s contemporaries, could it replicate the miracle that upended humanity’s understanding of space and time? This thought experiment, posed by Google DeepMind CEO Demis Hassabis, is sparking intense debate within the tech community.

This seemingly simple test actually draws a profound red line regarding the nature of intelligence: under conditions of complete information symmetry, can AI exhibit the creativity required to shift a scientific paradigm? Hassabis leans toward a negative conclusion — even if given all physics literature predating 1901, AI would still be unable to independently propose special relativity. The reason lies not in insufficient computing power or a lack of data, but in the fact that special relativity was never a simple act of knowledge completion. It fundamentally redefined how humanity understands time, space, and motion.

When discussing AI creativity, Hassabis draws a strict distinction between two capabilities: finding an optimal solution within an already-defined problem framework, versus proposing entirely new theories and hypotheses. In his assessment, only the latter truly touches the core of creativity. What this test really interrogates is this: if all the precursor materials for a scientific breakthrough are already in place, will a new theory naturally emerge from the soil of old knowledge, like ripened fruit?

To understand this question, we must return to a core concept in the philosophy of science — the “paradigm” introduced by Thomas Kuhn in The Structure of Scientific Revolutions. A paradigm can be understood as the set of problem-solving rules collectively accepted by a community of scientists in a given era. It dictates which questions are worth studying, what concepts should be used, what counts as valid evidence, and what kind of explanation is considered reasonable. Within the framework of Newtonian physics, time flowed uniformly, and space served as a fixed stage for the motion of objects. These elements were typically not treated as problems in themselves, but rather as fundamental premises scientists used to investigate other questions.

Within a stable paradigm, most scientific research is essentially problem-solving according to established rules. When conflicts arise between theory and experiment, scientists check whether measurements are accurate, whether parameters need adjustment, or whether a specific condition has been overlooked in the theory. As long as the basic framework remains functional, research continues to seek answers within the existing system. Kuhn termed this type of research “normal science.” Only when anomalies accumulate persistently and the existing framework proves incapable of providing a reasonable explanation, no matter what, can a scientific revolution truly erupt. A new theory redefines which questions are important and which concepts are valid, even fundamentally altering how scientists interpret the same set of experimental results. A paradigm shift is not about choosing a better answer among old ones; it is about changing the way we understand the problem itself.

The physics community at the end of the 19th century stood on the eve of just such a revolution. At that time, Newtonian mechanics could perfectly explain the macroscopic motion of objects, while Maxwell’s equations elegantly unified electricity and magnetism. Individually, both theoretical frameworks were extraordinarily successful; but when placed together, an unbridgeable chasm appeared. In Newtonian mechanics, velocities add directly. If a person on a moving train throws a ball forward, the speed of the ball measured by an observer on the platform should equal the train’s speed plus the ball’s speed relative to the train. By the same logic, observers in different states of motion measuring the speed of light should also obtain different results. Yet, the speed of light derived from Maxwell’s equations appeared to be a constant, unchanging value.

The subsequent Michelson-Morley experiment dealt an even heavier blow to traditional physics. If light must propagate through a mysterious medium called the “luminiferous aether,” and the Earth is moving at high speed through this aether, then the experiment should logically detect differences in the speed of light in different directions — the so-called “aether wind.” But the experimental result was negative; the speed of light was astonishingly consistent in all directions. Faced with this disruptive anomaly, the vast majority of scientists at the time still chose to patch up the old paradigm: perhaps the aether is dragged along by the Earth’s motion? Could objects physically contract when moving at high speeds? Is there some natural mechanism that makes the aether forever undetectable by humans?

Einstein faced the same set of contradictions but chose a radically different starting point. He stopped asking why the aether couldn’t be detected and ceased insisting that all observers must share a single absolute time. He boldly proceeded from two fundamental postulates: the laws of physics are the same in all inertial reference frames, and the speed of light in a vacuum is constant for all inertial observers. If both hold true simultaneously, then what needs adjustment is no longer a specific physical parameter, but humanity’s fundamental understanding of time and space. Whether two spatially separated events occur simultaneously no longer has an observer-independent absolute answer; time intervals and spatial lengths genuinely change depending on the observer’s state of motion. Einstein transformed “absolute time,” previously treated as an unquestionable given, into the core problem requiring re-examination.

This is precisely why the relativity test is so profoundly difficult. The relevant experimental data, theoretical contradictions, and mathematical tools all existed before 1901, but this body of knowledge would not automatically tell a researcher: perhaps the error lies not in some parameter, but in the way we understand time itself.

To explain why AI cannot invent relativity, one must also delve into the mechanisms by which current large language models (LLMs) acquire knowledge and generate what is perceived as “creativity.” LLMs, trained on vast corpora of text, learn the statistical relationships between words, concepts, and knowledge. The most fundamental training task is to predict the most likely next word based on existing context. When models reach sufficient scale, they master complex linguistic forms and exhibit emergent abilities in summarization, induction, reasoning, transfer, and combination. Faced with a clearly defined research problem, AI can rapidly read vast numbers of papers, organize different viewpoints, perform complex computational derivations, generate code, and even propose candidate hypotheses and experimental designs. A literature review that might take a research team months to complete can be outlined by AI in minutes.

However, AI’s judgment of “what is reasonable” is still primarily based on the distribution of existing knowledge within its training data. If the papers in its training corpus universally accept the existence of the aether, the model will treat the aether as a reasonable foundational premise; if most literature exhaustively investigates why the aether cannot be detected, it will continue searching for answers along that predetermined path. This does not mean AI completely lacks creativity. Its most common form of creation is the recombination of existing knowledge: articles, code, and hypotheses can all be entirely new, but the underlying materials and fundamental judgment criteria remain firmly rooted in the paradigm of old knowledge.

Furthermore, AI can find astonishing solutions within fixed rules that humans have never previously discovered. AlphaGo’s “Move 37” during its match against Lee Sedol is a classic example. That move was unprecedented in thousands of years of Go history, but the rules of Go, the objective of winning, and the feedback mechanisms were all pre-defined. What AI did was search a vast possibility space to find a new path untrodden by humans.

Relativity, however, demands a fundamentally different kind of creation. It did not seek a better solution within the established problem of “how to detect the aether,” but instead fundamentally re-examined whether the aether and absolute time needed to exist at all. It changed the question itself. AI can already find paths on an old map that humans have not walked, but relativity required the realization that the map itself might be wrong.

If one were to actually train a model with a knowledge cutoff of 1901, feeding it classical mechanics, Maxwell’s equations, and the relevant experimental results, could AI discover the contradictions within? Absolutely. It could clearly point out the conflict between the classical principle of velocity addition and the phenomenon of light-speed invariance, and systematically catalog the fatal difficulties the Michelson-Morley experiment posed for aether theory. Given sufficiently explicit prompts, it could even generate dozens of seemingly plausible explanations in one go. A discrepancy between theory and experiment can have many causes: measurement error, incorrect parameter settings, omission of a key condition, limited applicability of the theory, or a fundamental problem with the most basic underlying concepts. AI is extremely adept at listing these possibilities, but it struggles to independently judge at which level the problem truly lies.

Within the old paradigm, absolute time was never considered a hypothesis needing verification; it was more like a given, indisputable condition on an exam paper. Researchers would use this condition to solve problems, rarely questioning whether the premise itself was valid. Moving from “two theories are contradictory” to “absolute time might not exist at all” involves no step-by-step logical deduction chain. What this step requires is the overthrow of the very foundational premises used for deduction.

Current LLMs can already generate scientific hypotheses, but generating a hypothesis is still far removed from truly building a completely new theoretical system. If mainstream research in 1901 all revolved around the aether, AI would be more likely to propose a more sophisticated new aether model or design a more precise aether detection experiment. Even if it accidentally included “the aether might not exist” in a generated list of options, this would be just one unremarkable candidate among many. The real difficulty lies in theory selection: why abandon a fundamental premise widely accepted by the entire scientific community? Can this seemingly highly anomalous hypothesis unify more physical phenomena using fewer premises? In the absence of direct experimental evidence, why is it worth betting precious research resources to pursue it further?

Scientific breakthroughs require both generating new ideas and identifying, from a sea of possibilities, the directions truly worth betting on. AI can efficiently manufacture options but cannot independently make this kind of choice, which is laden with strong value judgments. Scientists keep probing, sometimes not because the equations don’t work out. A theory might already explain existing experiments, yet the entire system still feels insufficiently elegant, or noticeable cracks of inconsistency exist between different theories. It is this deep-seated dissatisfaction that drives researchers to continue searching for a more unified, more beautiful explanation. Currently, all of AI’s goals come from external commands. If a human asks it to explain an experiment, it seeks a solution to complete the task; if a human asks it to propose ten hypotheses, it generates ten. Once the task is finished, it does not continue to question, nor does it develop an intrinsic impulse for long-term commitment due to an incongruent worldview.

Even if AI were granted long-term memory, continuous tasks, and automated experimentation capabilities, the research goals and evaluation criteria would still be set by humans. Persistently executing an assigned problem is entirely different from proactively deciding which contradiction is worth a lifetime of inquiry or which premise needs to be completely overturned. AI can mimic the language of doubt with remarkable fidelity, but it lacks an existential need to think a problem through to its core. It can walk the paths within an old paradigm to their absolute limits, but it will not naturally generate the impulse that “this path itself might be wrong.”

Acknowledging this boundary of capability does not diminish AI’s revolutionary value for scientific research. Quite the opposite: a vast amount of future scientific work will be restructured by AI. Literature reading, data organization, computational derivation, code implementation, experimental design, and model screening — these repetitive, tedious, and extremely time-consuming parts can increasingly be handed over to AI. Literature conflicts that once took years to untangle could be precisely identified by AI in minutes; erroneous research directions that previously required extensive experimentation to eliminate can be filtered out at an accelerated pace through modeling. AI will also expose anomalies to researchers earlier, allowing scientists to dedicate more of their precious time to judgment, selection, and deep interpretation. Its significance for scientific research may far exceed the literature summaries and code generation we see today.

The more powerful AI becomes, the more human value in scientific research will shift upstream: deciding what to study, identifying which contradiction is truly worth pursuing, and judging when to decisively abandon an old explanatory framework. Hassabis’s relativity test reminds us that the most critical moment in scientific research sometimes lies not in answering questions faster, but in realizing that the question itself was asked incorrectly. Once a problem is clearly defined, AI can perform faster and more comprehensively than the vast majority of people. But when a breakthrough requires the researcher to doubt the most fundamental, most unquestionable premises, neither the volume of knowledge nor computational speed can automatically complete this breathtaking leap.

Thomas Edison once said, “Genius is one percent inspiration and ninety-nine percent perspiration.” The inspiration here is not a flash of insight appearing from nowhere; it originates from a human sensitivity to contradiction, a deep-seated doubt toward old premises, and the courage and ability to reconstruct an explanatory framework from a welter of phenomena when no standard answer exists. The material Einstein faced was no greater than that available to other top physicists of his era. AI can currently replace that 99% of perspiration, but it still cannot replace that decisive 1% of inspiration.