{"id":100250,"date":"2026-07-09T12:16:16","date_gmt":"2026-07-09T12:16:16","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/100250\/"},"modified":"2026-07-09T12:16:16","modified_gmt":"2026-07-09T12:16:16","slug":"i-built-a-self-improving-ai-and-so-can-you","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/100250\/","title":{"rendered":"I Built a Self-Improving AI, and So Can You"},"content":{"rendered":"<p>These days, the frontier AI labs are all racing to build <a href=\"https:\/\/www.wired.com\/story\/meta-earnings-superintelligence-q2-2025\/\" class=\"text link\" rel=\"nofollow noopener\" target=\"_blank\">self-improving models<\/a>. Some believe it\u2019s the surest route to <a href=\"https:\/\/www.wired.com\/story\/uncanny-valley-podcast-superintelligence\/\" class=\"text link\" rel=\"nofollow noopener\" target=\"_blank\">superintelligence<\/a>\u2014as AI improves itself in a mind-melting loop, the thinking goes, it will eventually surpass human comprehension (and perhaps even control).<\/p>\n<p class=\"paywall\">That\u2019s all well and good, but I have <a href=\"https:\/\/www.wired.com\/newsletter\/exclusive\/ai-lab\" class=\"text link\" rel=\"nofollow noopener\" target=\"_blank\">a newsletter<\/a> to produce. I wondered if recursive self-improvement might also be useful for me. Could I use AI to train and continually improve a model that automates some of this newsletter\u2019s busywork?<\/p>\n<p class=\"paywall\">After a week or so of experimenting, the answer appears to be a resounding\u2014and surprising\u2014hell yes. What\u2019s more, dabbling with self-improving models shows a different vision for how AI might unfold\u2014one that doesn\u2019t center on a handful of companies that control the whole industry.<\/p>\n<p>I started by trying out a simple self-improving loop<\/p>\n<p class=\"paywall\">To get my feet wet, I experimented with training a small language model from scratch\u2014by which I mean I dumped all the hard work on <a href=\"https:\/\/www.wired.com\/tag\/claude\/\" class=\"text link\" rel=\"nofollow noopener\" target=\"_blank\">Claude\u2019s<\/a> plate.<\/p>\n<p class=\"paywall\">I installed <a data-offer-url=\"https:\/\/github.com\/karpathy\/autoresearch\" class=\"external-link text link\" data-event-click=\"{&quot;element&quot;:&quot;ExternalLink&quot;,&quot;outgoingURL&quot;:&quot;https:\/\/github.com\/karpathy\/autoresearch&quot;}\" href=\"https:\/\/github.com\/karpathy\/autoresearch\" rel=\"nofollow noopener\" target=\"_blank\">AutoResearch<\/a>, which helps an off-the-shelf AI model build and improve a smaller model. AutoResearch is the brainchild of <a href=\"https:\/\/www.wired.com\/2015\/01\/karpathy\/\" class=\"text link\" rel=\"nofollow noopener\" target=\"_blank\">Andrej Karpathy<\/a>, a superstar AI researcher who helped found OpenAI, led AI work at Tesla, and recently <a data-offer-url=\"https:\/\/x.com\/karpathy\/status\/2056753169888334312?lang=en\" class=\"external-link text link\" data-event-click=\"{&quot;element&quot;:&quot;ExternalLink&quot;,&quot;outgoingURL&quot;:&quot;https:\/\/x.com\/karpathy\/status\/2056753169888334312?lang=en&quot;}\" href=\"https:\/\/x.com\/karpathy\/status\/2056753169888334312?lang=en\" rel=\"nofollow noopener\" target=\"_blank\">joined<\/a> Anthropic.<\/p>\n<p class=\"paywall\">I fired up Claude and gave it the recommended instruction: \u201cHi, have a look at program.md and let&#8217;s kick off a new experiment!\u201d While Claude did the hard stuff, I provided silicon (an <a href=\"https:\/\/www.wired.com\/tag\/nvidia\/\" class=\"text link\" rel=\"nofollow noopener\" target=\"_blank\">Nvidia<\/a> DGX, a desktop \u201csupercomputer\u201d designed for AI experimentation), the electricity (running hot for a few days straight), and a possibly ill-advised willingness to let the model skip all the usual permission checks in order to do its thing (let him cook!)<\/p>\n<p class=\"paywall\">I checked in on the AutoResearch project every few hours and marveled as Claude adjusted parameters and training regimes, looked at how this changed the smaller model\u2019s output, and went on refining it further.<\/p>\n<p class=\"paywall\">Here\u2019s what an early version of that smaller language model produced when I prompted it to complete the phrase \u201cIn the beginning \u2026\u201d<\/p>\n<p>\u201cIn the beginning of the beginning of the end of the end of the end end of end end end end end end end end beginning end end end end\u2026\u201d<\/p>\n<p class=\"paywall\">Not so brilliant. But later models, improved autonomously by Claude, got more coherent and less prone to insane, endless repetition. It\u2019s hardly GPT-5, but it showed a promising path toward continual improvement.<\/p>\n<p>My journey continued with something more complex\u2014and useful<\/p>\n<p class=\"paywall\">I already use an agent that relies on Claude to help me find noteworthy research papers, so I decided to see whether it was possible to build something that went beyond that.<\/p>\n<p class=\"paywall\">I turned to a tool from a startup called <a data-offer-url=\"https:\/\/www.primeintellect.ai\/\" class=\"external-link text link\" data-event-click=\"{&quot;element&quot;:&quot;ExternalLink&quot;,&quot;outgoingURL&quot;:&quot;https:\/\/www.primeintellect.ai\/&quot;}\" href=\"https:\/\/www.primeintellect.ai\/\" rel=\"nofollow noopener\" target=\"_blank\">Prime Intellect<\/a>, which uses AI to train a custom model for a specific task. I collected 100 or so previous \u201cElsewhere on the frontier of AI\u201d entries\u2014the bits and bobs of research that follow the main essay in <a href=\"https:\/\/www.wired.com\/newsletter\/exclusive\/ai-lab?sourceCode=CarveLeft\" class=\"text link\" rel=\"nofollow noopener\" target=\"_blank\">my newsletter<\/a>. Then, I created a Prime Intellect training environment and asked Claude to help me build my own model, which it dubbed Frontier_Paper_Curator, to find and summarize interesting papers.<\/p>\n<p class=\"paywall\">Claude found more papers and generated a bunch of synthetic data to help with training. It then tapped yet another model to assess Frontier_Paper_Curator\u2019s output, while the training environment also improved the model with reinforcement learning.<\/p>\n","protected":false},"excerpt":{"rendered":"These days, the frontier AI labs are all racing to build self-improving models. Some believe it\u2019s the surest&hellip;\n","protected":false},"author":2,"featured_media":100251,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[8],"tags":[179,1201,53,3154,25,580,182,157,52],"class_list":["post-100250","post","type-post","status-publish","format-standard","has-post-thumbnail","category-anthropic","tag-agentic-ai","tag-ai-lab","tag-anthropic","tag-anthropic-claude","tag-artificial-intelligence","tag-chatgpt","tag-claude","tag-openai","tag-research"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/100250","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=100250"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/100250\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/100251"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=100250"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=100250"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=100250"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}