{"id":31023,"date":"2026-05-07T14:17:33","date_gmt":"2026-05-07T14:17:33","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/31023\/"},"modified":"2026-05-07T14:17:33","modified_gmt":"2026-05-07T14:17:33","slug":"transformer-architecture-superpowers-and-the-march-toward-agi-2","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/31023\/","title":{"rendered":"Transformer Architecture, Superpowers, And The March Toward AGI"},"content":{"rendered":"<p><img decoding=\"async\" class=\" top-image\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/05\/0x0.jpg\" alt=\"CPU. AI - Artificial Intelligence and machine learning concept\" data-height=\"977\" data-width=\"1738\" fetchpriority=\"high\" style=\"position:absolute;top:0\"\/><\/p>\n<p>digital transformation. AI data. innovations and technology.<\/p>\n<p>getty<\/p>\n<p>When you scratch the surface of some of the best new development and research on AI, you see that we\u2019re likely very close to artificial general intelligence, that next generation or step that will give AI super-human powers of thought.<\/p>\n<p>I watched this<a href=\"https:\/\/www.youtube.com\/watch?v=XeTuLyOBY_0&amp;t=353s\" target=\"_blank\" rel=\"nofollow noopener noreferrer\" data-ga-track=\"ExternalLink:https:\/\/www.youtube.com\/watch?v=XeTuLyOBY_0&amp;t=353s\" aria-label=\"entire video where Sam Altman talks\"> entire video where Sam Altman talks<\/a> about novel architectures, and what they mean for the industry, about replacing the transformer as an attention mechanism, and the result of that approach.<\/p>\n<p>It\u2019s wild stuff, and it\u2019s a big deal.<\/p>\n<p>Using methods like compressed context to build new architectures means the AI entities will better understand the world around them, which in turn will translate into much more capability in all kinds of tasks. Altman used the example of Apple\u2019s LITO model which can compile a three-dimensional object from one two-dimensional image.<\/p>\n<p>But there\u2019s much more to the story.<\/p>\n<p>The Transformer and its Legacy<\/p>\n<p>What about the transformer? Is it so central that it\u2019s hard to replace?<\/p>\n<p>Basically, as Altman and others explain, the transformer replaced more primitive architectures with something that focuses AI\u2019s attention, for maximum efficiency and more targeted convergence on a given problem. But now, approaches like subquadratic scaling and liquid networks are making the transformer effectively obsolete, at least in some ways.<\/p>\n<p>In a segment at our Imagination in Action event April 9\u201310 (April\u2019s IIA event is an annual conference that I help to facilitate here at MIT), my colleague from Link Ventures Dave Blundin interviewed Peter Danenberg of Google, and Alexander Amini of Liquid AI (further disclaimer, I have also been involved in Liquid AI).<\/p>\n<p>Amini mentioned how big players like Alibaba and Qwen are moving away from, or beyond, the transformer.<\/p>\n<p>\u201cWe&#8217;re seeing this shift already today, that you know, the top models, even trillion parameter models now, are hybrid models between the transformers and other models as well,\u201d he said.<\/p>\n<p>However, Amini\u2019s evaluation was nuanced, and he suggested that we may still use transformers for some specific kinds of projects.<\/p>\n<p>\u201cEvery architecture is good for its own situation, right?\u201d Amini said. \u201cTransformers are good for some things, but for some hardware, for some use cases, for some design decisions, there are much better architectures, mathematical operators that exist.\u201d<\/p>\n<p>\u201cIt may be the case that transformers are slightly saturated,\u201d added Danenberg. \u201cI don&#8217;t know if we&#8217;re going to get the next step function from transformers.\u201d<\/p>\n<p>Google and the TPU<\/p>\n<p>The panel also had a discussion about how Google\u2019s tensor processing hardware is competing with the Nvidia designs that drove the latter company to the top of the heap on the American stock market.<\/p>\n<p>\u201cOne of the things that that&#8217;s been evolving very quickly is somebody will have a brilliant algorithmic breakthrough, for example, and then within the Google empire, immediately, there&#8217;s a conversation on how to convert it to silicon,\u201d Blundin said.<\/p>\n<p>That, Amini suggested, makes sense in the context of a market where hardware and digital operations should be, in his words, \u201cmarried\u201d together.<\/p>\n<p>\u201cThis is a philosophy that we have at liquid, that architecture and hardware should be married together, right?\u201d Amini said. \u201cAnd that if you want to build the best quality, AI, quality is also dependent on speed efficiency, energy efficiency as well. So you need these things to be co-optimized together.\u201d<\/p>\n<p>Different Kinds of Customers<\/p>\n<p>Amini suggested there\u2019s a big difference between general-purpose use cases, and those specific to an industry, where a client wants something very targeted.<\/p>\n<p>\u201cIt comes down to value delivery at the end of the day, and delivering value for different types of use cases,\u201d he said. \u201cIf you want to deploy, let&#8217;s say, an AI to be a conversational assistant to you in the car, you probably don&#8217;t need that AI to be optimized on PhD-level physics, right? There&#8217;s a different type of distribution that will be important for those domains, and those enterprises, that is fundamentally different in scale than the big AI providers, and that will yield itself in the forms of new, efficient architectures that effectively get distilled down from the big ones.\u201d<\/p>\n<p>Danenberg painted a picture of competitive business ecosystems.  <\/p>\n<p>\u201cStartups and a bunch of businesses run constellations of small models, almost like the Unix philosophy, each one of which sort of does one thing, and does it well,\u201d he said. \u201cAnd really, the job of the business, in that sense, is just to sort of orchestrate these tiny models.\u201d<\/p>\n<p>Understanding the \u201cBeast\u201d<\/p>\n<p>Later in the discussion, Danenberg explained how, at the outset of debate about AGI, he struggled to figure out what people would practically use the technology for. Then, he said, he started to imagine how this sort of thing would emerge.<\/p>\n<p>\u201cWhen this really slow, sort of expensive, big model with a lot of world knowledge baked in, when that thing is following you around all day, it turns out that it comes up with these really interesting connections that you may not have anticipated,\u201d he said. \u201cIt may be that there are certain cases where you actually do want these, these \u2018beasts.\u2019\u201d<\/p>\n<p>Blundin added his two cents.<\/p>\n<p>\u201cThese things, they cross the uncanny valley very, very quickly,\u201d he said, \u201cand they become something you just like, can&#8217;t live without. And you know, when they&#8217;re just below that level, they&#8217;re icky as all hell, but when they cross the line, \u2018whoa,\u2019 instantaneously.\u201d<\/p>\n<p>Change Management<\/p>\n<p>There was more in the discussion, about Moore\u2019s law, specialists vs. generalists, and the appeal of general purpose Nvidia hardware, to name a few.<\/p>\n<p>\u201cI feel like the world in AI will move towards hierarchical memory more and more,\u201d Amini offered. \u201cAnd right now, AI models don&#8217;t have this right &#8211; we have moved entirely away from hierarchical memory, because all of the memory is stored in the context, and you basically invest into storing all of that as tokens.\u201d<\/p>\n<p>Blundin, toward the end, asked about what jobs could be \u201c10x-ed\u201d or made ten times more effective with AI.<\/p>\n<p>\u201cI&#8217;m actually at a loss sometimes to think of things that aren&#8217;t 10x-able in that sense,\u201d Danenberg said.  \u201cSo I don&#8217;t know, maybe we&#8217;re shooting for 95% to 99%.\u201d<\/p>\n<p>\u201cI think it&#8217;s the majority today,\u201d Amini added.<\/p>\n<p>The Big Question<\/p>\n<p>If we really can 10x 99% of human jobs, or replace that much productivity, what happens to human workers? That\u2019s a big question that\u2019s making the rounds right now, in 2026, as we see new research on more powerful models continue. Stay tuned.<\/p>\n","protected":false},"excerpt":{"rendered":"digital transformation. AI data. innovations and technology. getty When you scratch the surface of some of the best&hellip;\n","protected":false},"author":2,"featured_media":31024,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[4],"tags":[6744,3013,20034,20033,17486,20032],"class_list":["post-31023","post","type-post","status-publish","format-standard","has-post-thumbnail","category-agi","tag-agi","tag-artificial-general-intelligence","tag-danenberg-amini","tag-liquid-ai","tag-peter-danenberg","tag-transformer-architecture"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/31023","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=31023"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/31023\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/31024"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=31023"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=31023"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=31023"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}