At a crossroads where the AI industry is being pushed toward commercial deployment by large language models and foundation models, Richard Sutton, the 2024 Turing Award laureate and “Father of Reinforcement Learning,” has dropped a bombshell. He bluntly stated that the current industry over-relies on human data, exhibiting significant inefficiencies in energy and computational resource utilization. True superintelligence, he argued, must completely shed human prior biases and evolve autonomously from its own experience, much like an infant.

On July 19, at the “Thinkers Forum” of the 2026 World Artificial Intelligence Conference (WAIC), Sutton delivered a keynote titled “Cultivating Superintelligence from Experience via Reinforcement Learning” and subsequently sat for interviews with media outlets including China Business News. In a nearly 40-minute dialogue, the newly minted Turing Award winner not only systematically critiqued the pitfalls of current mainstream technical approaches but also, for the first time, comprehensively introduced his latest theoretical framework, “OaK” (Options and Knowledge), while revealing his dual-track strategy spanning entrepreneurship and academic research.

Farewell to “Human Data”: Sutton’s “Age of Experience” Manifesto

Sutton expressed disappointment with the trajectory of the AI industry over the past eight years. He recalled that during the AlphaGo and AlphaZero era, the industry had already validated the immense potential of self-learning through simulated experience. However, subsequently, the entire sector was swept into a dividend period of relying on human data, neglecting the fundamental path of learning from experience.

“Looking back at the industry’s trajectory over the past eight years, I feel a certain degree of disappointment,” Sutton stated candidly in the interview. “This period witnessed a wave of excessive focus on ‘human data.’ This deviation from the path is regrettable.”

During his speech, Sutton reiterated the thesis of his famous AI research essay, The Bitter Lesson. He pointed out that AI researchers always want to hard-code “human knowledge” into AI, but in the long run, this approach actually hinders AI’s evolution. Sutton emphasized: “We should seek the underlying first principles of intelligence, rather than being misled by the specific details generated by our own minds in a particular world. We should not try to build in the patterns that humans have already discovered; instead, we should let the agent discover them on its own.”

To achieve true general intelligence, Sutton set a highly challenging goal: building a trillion-parameter model capable of real-time learning and planning using only about 20 watts of power. He noted that the human brain requires only about 60 watts to handle extremely complex tasks, proving that there is enormous room for optimization in digital computation.

“Traditional digital computing is constrained by the von Neumann architecture, where data frequently migrates between storage and processors, leading to extremely low energy efficiency,” Sutton explained. He is working to build a more biologically-inspired, brain-like computing strategy, centered on more thorough distribution and parallelization. Based on this, future agents will no longer rely on large volumes of reproducible data for learning but will dynamically adjust their behavior through autonomous shaping in brief scenarios with limited samples.

Sutton predicted that the industry is accelerating toward an “Age of Experience.” If the timeline is extended to eight years, he firmly believes this path will gain as much industry recognition as LLMs currently enjoy. He revealed that several leading companies focused on experiential learning have already been established, including Yann LeCun’s company and his own Openmind.

Critiquing “Low-Level Physics Simulators” and Proposing the OaK Architecture

On the theoretical front, Sutton sharply criticized the current industry’s definition of “world models.” He considers the approach of turning world models into “low-level physics simulators” an “extremely narrow” understanding.

“In my theoretical system, I would absolutely not call a simple neural network a ‘model’ — that is an abuse of the term,” Sutton explained in his speech. “I reserve the word ‘model’ exclusively for transition models, which should encompass all of humanity’s high-level knowledge about the world. For example, ‘If I take this job, will I be happy?’ — that is the kind of question a model should answer. Human transition models should not be as trivial as low-level physics; they should carry high-order world knowledge.”

This implies that models need to master “abstraction” — the ability to construct increasingly grand perceptions of actions and states. To achieve this high-level abstraction, Sutton proposed the novel OaK architecture. Its core lies in enabling AI to achieve “leap-style” macro-planning through temporal abstraction (options) and state abstraction (knowledge).

Within the OaK architecture, for every state feature relevant to human life, the agent must learn a skill to “control and achieve that feature.” The core mechanism by which the agent constructs its own mind is continuously posing sub-problems and solving them through “options.” Sutton summarized this as “completely handing over the power to define sub-problems to the agent.”

Sutton used infant behavior as an analogy for this process. Infants exhibit an adaptive mechanism that drives their next decisions based on their own cognitive level. When autonomously deconstructing and attempting to solve a series of low-level sub-problems, they dynamically evaluate their marginal learning effect. Once they find that something no longer provides new cognitive input and generates a sense of “boredom,” they actively switch to the next unsolved problem.

This mechanism ensures the maximization of continuous learning efficiency. Sutton’s core proposition is that “learning brings positive feedback.” If we can endow agents with the ability to measure their own learning progress and convert it into an internal implicit reward mechanism, that would be a profoundly transformative architectural design. He specifically noted that this also aligns with the ancient Chinese cognitive philosophy of “Is it not a pleasure to learn and constantly practice what one has learned?”

Although the OaK architecture can theoretically achieve a closed loop, Sutton admitted that it cannot be realized immediately. The arrival of artificial general intelligence may still take 5, 10, or even 20 years.

Dual-Role Entrepreneurship: For-Profit and Non-Profit in Parallel

Just this week, Sutton announced the co-founding of a new company, Oak Lab, with University of Alberta doctoral researcher Khurram Javed. The venture aims to break away from current deep learning paradigms and build AGI with entirely new concepts, developing AI agents capable of learning independently and continuously through their own first-person experience.

Explaining why he chose to start a company now, Sutton said that Oak Lab is a for-profit entity, but it complements his previous non-profit activities at the Openmind Research Institute. He had previously participated in another for-profit company focused on core technologies for three years, and that positive experience made him realize that establishing an independent organization allows for greater focus in research direction.

“Although the operating models differ, the core vision is highly consistent: ‘understanding how intelligence works, and then sharing it with the world,'” Sutton emphasized. He insisted on global open-sourcing because he believes the scope of knowledge dissemination is positively correlated with the effect of technology democratization. Increasing the transparency of cutting-edge knowledge can effectively reduce the potential risk of technology misuse.

However, Sutton also acknowledged that frontier research requires substantial resource support, which is precisely where the value of a for-profit company lies. Certain core technical information needs short-term protection as patents. While no technical secret can be permanently locked down in the long run, commercializing valuable innovative ideas in the short term is an inevitable path to continuously reinvest in scientific research.

Sutton is currently raising funds for Oak Lab, and progress has been quite smooth. Although relevant partnerships have yet to be finalized, he believes the capital and industry thirst for “experiential learning” is already evident. Following the explosion of LLMs and foundation models, the market is eagerly searching for the next breakthrough direction.

Embodied Intelligence and Social Necessity

In the field of embodied intelligence, Sutton also revealed concrete practical progress. The “Robot Kindergarten” project he participates in is a collaboration between Openmind and Tashan Technology in Beijing, aimed at building a dedicated experimental space for robots.

The project seeks to solve the pain point of traditional robots being prone to damage because they cannot tolerate trial-and-error. Traditional robots are designed from the outset to precisely execute predetermined commands, not to explore the unknown through experimentation, which often results in demonstrations lacking long-term viability. The core of the system developed by Sutton’s team lies in creating robotic systems capable of withstanding prolonged, high-intensity physical trial-and-error, with exploration cycles extending from days to months.

“This is similar to how human infants spend years conquering foundational skills like walking and object grasping,” Sutton said. Once robots can comprehensively accumulate experience through reinforcement learning and trial-and-error, the robotics industry is bound to undergo a disruptive transformation.

Regarding whether future robots will solve problems like elderly care and teacher shortages, Sutton confirmed that this is already incorporated into the industry’s strategic blueprint. Robots will play a critical role in fields with high labor demand, such as elder care. Especially against the backdrop of a potential future decline in the labor force dividend, AI filling labor gaps will become an inevitable trend.

The vision of “Robot Kindergarten” lies in constructing a new type of socialized symbiotic relationship: treating robots as individuals with growth potential, nurturing them carefully to mature their minds, and ultimately guiding them to integrate into and serve human society.

Advice for Young People: Reject Blind Conformity, Persist in Writing

At the end of the interview, Sutton offered advice to the younger generation of researchers and young people deeply affected by AI.

He stressed that “intelligence” and “computation” should be clearly distinguished at the foundational logic level. Computational dividends belong to engineering extension, while the arrival of general intelligence represents a dimensional leap of disruption. The industry currently estimates the median year for this technological singularity to arrive as 2040, but the technical uncertainty around that estimate remains extremely high.

Faced with the cyclical law that “the more things change, the more the essence stays the same,” Sutton believes the most valuable self-investment for young people is to build and gradually expand a cognitive fortress that they can deeply deconstruct, persisting in independent thinking based on first principles.

“As a young person, you should not blindly follow authority. Insist on autonomous verification through quantifiable methods, and refuse to accept ready-made answers that are directly instilled,” Sutton stated solemnly. “There are no absolute authorities in the scientific community — even when speaking with so-called ‘authoritative status,’ I must reiterate this point. Everyone is the ultimate person responsible for their own cognitive conclusions and must possess the critical thinking skills for independent evaluation.”

Additionally, he provided a concrete methodology: maintain a daily writing practice. For example, write one page of notes each day, systematically organizing, questioning, and reshaping your own thinking. This requires high-intensity mental discipline and time accumulation, typically needing more than a year of persistence before the compounding effect becomes apparent. Sutton said: “If you hope your views will earn the respect of others, the prerequisite is that you must invest sufficient sincerity in your own thinking, commit it to paper, and optimize your logical expression through textual refinement.”

The ideas of the Father of Reinforcement Learning are moving from theory to practice. Whether through the proposal of the OaK architecture or the dual-track operation of Oak Lab and Openmind, Sutton is attempting to find an evolutionary path for artificial intelligence that breaks away from human prior knowledge and returns to the essence of experience. At a moment when capital and industry are anxiously seeking the next breakthrough direction beyond LLMs, Sutton’s “Age of Experience” may just be dawning.