{"id":900395,"date":"2026-06-29T08:35:15","date_gmt":"2026-06-29T08:35:15","guid":{"rendered":"https:\/\/www.europesays.com\/us\/900395\/"},"modified":"2026-06-29T08:35:15","modified_gmt":"2026-06-29T08:35:15","slug":"can-robots-fail-forward-teaching-ai-to-learn-from-its-mistakes","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/us\/900395\/","title":{"rendered":"Can robots \u201cfail forward\u201d? Teaching AI to learn from its mistakes"},"content":{"rendered":"<p><strong>Editor\u2019s Note:<\/strong> This work relates to Department of Navy award N000142212474 issued by the Office of Naval Research.<\/p>\n<p>For years, the gold standard for training artificial intelligence has been to show it how to succeed. Whether it was Google DeepMind\u2019s AlphaGo learning from millions of human professional moves or a robot being trained on motion-tracking data from humans in a maze, the assumption was that AI needs an expert model to mimic.<\/p>\n<p>However, <a href=\"https:\/\/klesse.utsa.edu\/faculty\/profiles\/cao-yongcan.html\" target=\"_blank\" rel=\"noopener nofollow\"><strong>Yongcan Cao<\/strong><\/a>, a researcher at The University of Texas at San Antonio, is flipping that logic on its head. He is developing a new way for autonomous systems to learn by focusing not on what goes right, but on what goes wrong.<\/p>\n<p>\u201cHumans take risks and learn from failure,\u201d said Cao, PhD, who holds the Mary Lou Clarke Endowed Distinguished Professorship in the Margie and Bill Klesse College of Engineering and Integrated Design. \u201cThink about a baby learning to walk. They stand up, they fall down, and they learn from that fall. We are looking at how to give that same mechanism to AI.\u201d<\/p>\n<p><strong>Solving the \u201csparse reward\u201d problem<\/strong><img fetchpriority=\"high\" decoding=\"async\" class=\" wp-image-28845\" src=\"https:\/\/www.europesays.com\/us\/wp-content\/uploads\/2026\/06\/cao-yongcan.jpg\" alt=\"Portrait of Yongcan Cao\" width=\"228\" height=\"285\"\/>Yongcan Cao<\/p>\n<p>In the world of reinforcement learning \u2014 a type of machine learning where an AI learns through trial and error \u2014 researchers often face the \u201csparse reward\u201d problem. In many complex tasks, an AI agent only receives a \u201creward\u201d or feedback when it successfully completes a task. If the task is highly complex, the agent might wander aimlessly for millions of attempts without ever stumbling upon a success, meaning it receives no data to learn from. Success becomes a matter of chance, much like the \u201cinfinite monkey theorem,\u201d where a monkey idly typing on a typewriter may eventually write something coherent, but it\u2019s unlikely to happen any time soon.<\/p>\n<p>To bridge this gap, engineers traditionally use \u201cexpert demonstrations,\u201d where a human or a pre-programmed system shows the AI exactly how to perform the task. While effective, this data is expensive, difficult to collect and sometimes impossible to obtain for training in new or dangerous environments.<\/p>\n<p>Cao\u2019s solution, detailed in recent publications including an award-winning abstract for the 2025 International Conference on Autonomous Agents and Multiagent Systems (AAMAS), is a framework called On-Policy Reinforcement Learning from Failure, or \u201cOn-F.\u201d<\/p>\n<p><strong>How On-F differs from standard algorithms<\/strong><\/p>\n<p>Standard reinforcement learning (RL) is akin to a student taking a final exam without ever receiving feedback on their homework throughout the course; they only find out if they passed or failed at the very end.<\/p>\n<p>The On-F framework introduces a \u201cdiscriminator,\u201d which acts as a judge. Instead of waiting for a final success, the system constantly compares the AI\u2019s current actions to a database of known failures \u2014 data that is \u201ccheap and abundant,\u201d according to Cao.<\/p>\n<p>Through a process called \u201creward densification,\u201d the AI receives constant, incremental feedback that encourages it to attempt the task in new ways. If its current path looks too much like a previous failure, the discriminator provides a \u201cpenalty.\u201d This pushes the AI to try something different, effectively \u201cdensifying\u201d the feedback loop so the agent is always learning, even when it hasn\u2019t succeeded yet.<\/p>\n<p>\u201cIf you imagine a drone flying a specific flight path and failing to locate a target, you don\u2019t want the drone to retrace the same route or fly just a few feet to the left or right. You\u2019d want the drone to try a significantly new approach, such as changing altitude or switching to a wide-angle view,\u201d Cao explained.<\/p>\n<p>When failure is systematic in this way, \u201cfailure alone can be used to learn desirable actions even if we don\u2019t have an expert model,\u201d Cao said. However, he also noted that when models are supplied with a mix of learning from failure and learning from demonstration, \u201cwe get even better outcomes.\u201d<\/p>\n<p><strong>Competitive performance and real-world results<\/strong><\/p>\n<p>The findings suggest that this \u201cfail-forward\u201d approach is more than just a theoretical concept. In simulated environments where digital robots were tasked with learning to stand or walk, the AI models using the On-F framework performed as well as \u2014 and in some cases better than \u2014 models trained using expensive expert data.<\/p>\n<p>The framework was validated using the Gymnasium simulation suite, specifically on tasks like the \u201cPointMaze,\u201d where an agent must navigate through a labyrinth and reach a target with very limited feedback.<\/p>\n<p>The research is supported by a $502,051 grant from the Office of Naval Research (ONR) as part of a multi-year project aimed at making autonomous systems, such as drones, as fast and efficient at decision-making as humans.<\/p>\n<p>\u201cAs humans, we don\u2019t take information at face value; we think critically,\u201d Cao said. \u201cFrom a robotics side, if a drone is pursuing a target and one path is risky, the system needs to analyze that risk based on what it knows could go wrong.\u201d<\/p>\n<p><strong>How On-F stands to shape industry <\/strong><\/p>\n<p>The ability to train AI using failure data could have massive implications across several industries. In manufacturing and robotics, it could drastically lower the cost of training new systems because engineers would no longer need to spend hundreds of hours creating \u201cperfect\u201d training scenarios.<\/p>\n<p>In the realm of autonomous vehicles and drones, the technology could lead to more robust navigation systems that are better at avoiding obstacles by recognizing the \u201csignature\u201d of a potential collision before it happens.<\/p>\n<p>Much like how AlphaGo began playing the ancient Chinese game of Go in ways humans had never imagined, Cao hopes that his AI models will also clear hurdles humans have yet to overcome \u2014 such as designing seemingly impossible surgeries or manufacturing first-of-its-kind nanotechnology.<\/p>\n<p>\u201cMy goal is always to help get the word out and develop more actionable insights,\u201d Cao said. \u201cPersonally, I feel this is a very helpful way to enable autonomous systems to learn as efficiently as humans do.\u201d<\/p>\n<p>As Cao\u2019s work with the ONR continues through July 2026, his team at UT San Antonio is exploring next steps, including how to refine these \u201cjudging\u201d systems to handle more subjective failures, potentially opening the door for AI that can assist in healthcare, logistics and complex disaster-response scenarios where there is no manual for success, only a history of mistakes to avoid.<\/p>\n<p>Any opinions, findings, and conclusions or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the views of the Office of Naval Research.<\/p>\n","protected":false},"excerpt":{"rendered":"Editor\u2019s Note: This work relates to Department of Navy award N000142212474 issued by the Office of Naval Research.&hellip;\n","protected":false},"author":3,"featured_media":900396,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":"","_share_on_mastodon":"0"},"categories":[5133],"tags":[5229,7202,7203,358,3187,67,586,132,5230,68,2969],"class_list":["post-900395","post","type-post","status-publish","format-standard","has-post-thumbnail","category-san-antonio","tag-america","tag-san-antonio","tag-sanantonio","tag-texas","tag-tx","tag-united-states","tag-united-states-of-america","tag-unitedstates","tag-unitedstatesofamerica","tag-us","tag-usa"],"share_on_mastodon":{"url":"","error":""},"_links":{"self":[{"href":"https:\/\/www.europesays.com\/us\/wp-json\/wp\/v2\/posts\/900395","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/us\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/us\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/us\/wp-json\/wp\/v2\/users\/3"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/us\/wp-json\/wp\/v2\/comments?post=900395"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/us\/wp-json\/wp\/v2\/posts\/900395\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/us\/wp-json\/wp\/v2\/media\/900396"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/us\/wp-json\/wp\/v2\/media?parent=900395"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/us\/wp-json\/wp\/v2\/categories?post=900395"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/us\/wp-json\/wp\/v2\/tags?post=900395"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}