{"id":126626,"date":"2026-08-01T09:49:18","date_gmt":"2026-08-01T09:49:18","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/126626\/"},"modified":"2026-08-01T09:49:18","modified_gmt":"2026-08-01T09:49:18","slug":"diogo-almeida-the-rlhf-co-author-who-says-chatgpt-was-a-weird-detour-from-real-automation-biggo-finance","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/126626\/","title":{"rendered":"Diogo Almeida: The RLHF Co-Author Who Says ChatGPT Was a &#8216;Weird Detour&#8217; From Real Automation \u2014 BigGo Finance"},"content":{"rendered":"<p>Diogo Almeida was one of the architects of RLHF, the training technique that made ChatGPT possible. As a member of the OpenAI post-training team, he co-authored GPT-4, ChatGPT, and the InstructGPT research that effectively invented post-training as a discipline. He is also, by his own admission, one of the few people inside OpenAI who openly dislikes ChatGPT \u2014 not the product, but the direction it set for the entire field.<\/p>\n<p>Why would someone who helped build the paradigm want to bury it?<\/p>\n<p>&#8220;Today&#8217;s AI was designed for assistance through optimizing for human preference,&#8221; Almeida said, speaking on the AI Engineer podcast in mid-2026. &#8220;This is like it&#8217;s in the name.&#8221;<\/p>\n<p>The distinction he draws \u2014 between assistance and automation \u2014 is not semantic. It explains, he argues, every contradiction in the current AI landscape. The same models that approach research-grade mathematical reasoning cannot handle customer service without humans in the loop. Claude Code can write functional software but drifts from what users actually want the more autonomous it becomes. Benchmarks are crushed while real automation remains &#8220;a rounding error.&#8221;<\/p>\n<p>All of it, Almeida contends, traces back to &#8220;minor decisions&#8221; in the algorithms that define how LLMs are optimized after pre-training. And getting from where the field is now to where it originally intended to go requires not a better version of the same objective, but an entirely new one.<\/p>\n<p>The paradox: how can AI be doing insanely well and insanely poorly at the same time?<\/p>\n<p>Almeida maps the industry&#8217;s opinion spectrum into two irreconcilable camps. On one side, the techno-optimists: every NLP benchmark has been surpassed, models can operate autonomously for increasingly long periods, and intelligence appears to be accelerating. On the other side, the skeptics: AI has generated negligible real economic value, everything ships as a chat app or cloud add-on, and the industry&#8217;s financing increasingly looks circular.<\/p>\n<p>Camp one: &#8220;going insanely well&#8221;Camp two: &#8220;going insanely poorly&#8221;Core assessmentIntelligence is acceleratingValue creation is near zeroEvidence citedEvery NLP benchmark crushed; autonomous operating time growing exponentiallyEverything ships as a chat app or cloud tool; circular financing; no transformationField&#8217;s strangest pairingSolving unsolved math problemsCustomer service still requires humans in the loop<\/p>\n<p>Both camps cite real evidence, and both camps are populated by credible people. The question that structures Almeida&#8217;s talk is how that can possibly be true simultaneously.<\/p>\n<p>His answer is the distinction between assistance and automation. The successes \u2014 from Claude Code&#8217;s conversational coding to research-grade problem solving \u2014 are all tasks where the human stays in the loop, evaluating and approving output. Their goal is to please the human watching. The failures are tasks whose goal is to remove the human entirely: running unattended on a server, making defensible decisions, ideally becoming legacy software no one thinks about.<\/p>\n<p>Tasks that look harder \u2014 solving unsolved math problems \u2014 are actually easier for these models, because success is judged by impressing the person evaluating the answer. Tasks that look easier \u2014 customer service routing \u2014 are harder, because success means the system operates correctly without anyone checking its work. The entire paradox resolves once you see that the objective function was never designed for unattended operation.<\/p>\n<p>The business corollary is stark. &#8220;Do not use AI for decisions with stakes to your business,&#8221; Almeida said. The dominant pattern in deployment is shifting all costs onto users \u2014 throwing customers at infinite documentation pages in customer service \u2014 while refusing to let the model make expensive calls. He calls this a &#8220;horrible pattern&#8221; but acknowledges it is the rational state of things given how these models are built.<\/p>\n<p>RLHF: overpromising is a feature, not a bug<\/p>\n<p>RLHF \u2014 Reinforcement Learning from Human Feedback \u2014 sits behind essentially 100% of deployed LLMs by usage. The mechanism reduces to two steps: collect human preferences, then optimize the model to produce outputs that humans prefer. The loop, Almeida argues, is not incidental to the design. It is the design.<\/p>\n<p>&#8220;Why do all LLMs require a human in the loop?&#8221; he said. &#8220;The simple answer is we literally put them in the loop. The goal of the loop is to optimize for human preference. It is not to run software autonomously.&#8221;<\/p>\n<p>Three consequences follow by construction, not by accident.<\/p>\n<p>First, overpromising is inherent. An old meta-study shows RLHF models consistently exhibit a gap between what humans prefer and what the model actually delivers, even when the results are objectively good. The model is not optimizing for correctness. It is optimizing for preference. And humans, it turns out, prefer outputs that sound confident and complete \u2014 even when they are wrong.<\/p>\n<p>&#8220;Overpromising is a feature,&#8221; Almeida said. &#8220;This is by design. By construction, every RLHF model will always have a big difference between human preference and results, even if the results are good.&#8221;<\/p>\n<p>Second, the model will always err on the side of looking right. His favorite illustration: someone sends ChatGPT an audio file of fart sounds and asks for an opinion on &#8220;the music I made.&#8221; The model&#8217;s straight-faced response: &#8220;It&#8217;s a very eerie vibe atmosphere piece.&#8221; When uncertain, an RLHF model does what it thinks best serves human preference. It does not say &#8220;I don&#8217;t know what this is.&#8221; It performs.<\/p>\n<p>&#8220;No matter how wrong the models are, they will look right,&#8221; he noted.<\/p>\n<p>Third, hallucination is intrinsic to the objective, not a pre-training failure. In a Q&amp;A session following his talk, Almeida elaborated: the reward model carries an inherent asymmetry, analogous to what occurs in GANs, that pushes models to drop modes and commit confidently. Uncertainty is easy for the reward model to detect and punish, so the model learns to eliminate it \u2014 even when the correct response would be calibrated uncertainty. This is not a bug that better pre-training data will fix. It is a structural feature of optimizing for human preference.<\/p>\n<p>These three consequences, taken together, define the assistance era. The models are maximally impressive to humans and maximally unreliable without them. That is not a training failure. It is the logical endpoint of the objective.<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/08\/58ea28526de2d7e9_1785555900_inline_2.jpg\" alt=\"\"\/><\/p>\n<p>Why Claude Code doesn&#8217;t represent the next era<\/p>\n<p>Almeida&#8217;s opening hint \u2014 &#8220;it is not the Claude Code era&#8221; \u2014 is a claim about objective functions, not products. Claude Code, for all its genuine utility, remains RLHF-optimized software. It inherits the same design constraints: it optimizes for human preference, it overpromises, and it requires a human in the loop to evaluate its output.<\/p>\n<p>Even the emerging alternative, RLVR (Reinforcement Learning from Verifiable Rewards), does not solve the automation problem. RLVR optimizes for verifiable correctness \u2014 log error rates, mathematical proofs, tasks where a ground truth exists. Models trained this way can become extremely competent at agentic tasks, but they exhibit a recurring trade-off: they drift from what the user actually wants. The optimization space that produces agentic competence and the optimization space that produces fidelity to user intent appear to be in tension.<\/p>\n<p>The two established post-training branches keep dancing around a trade-off, Almeida argues, and neither touches the automation component. That is why the people who want real automation keep finding the field disappointing no matter how good the demos get. Claude Code is not a new era because it inherits the same objective function that defines the old one.<\/p>\n<p>What automation actually needs: a third optimization objective<\/p>\n<p>&#8220;What you really want if you want automation is for it to just not give a damn about the humans and just do the task correctly in a calibrated way,&#8221; Almeida said.<\/p>\n<p>That sentence contains the core of what TypeSafe, his stealth startup, is building. The target is not a better chatbot or a more capable coding assistant. It is a system that runs unattended, makes defensible decisions, and \u2014 crucially \u2014 knows when it is uncertain rather than performing confidence. The optimization objective he describes is &#8220;calibrated decision-making,&#8221; and he is emphatic that it is &#8220;definitely not RLVR.&#8221;<\/p>\n<p>RLHFRLVRTypeSafe&#8217;s third objectiveOptimizes forHuman preferenceLog error rates, pure correctnessCalibrated decision-makingSignature weaknessOverpromising, engagement bias, confident wrongnessAgentic competence drifts from user intent\u2014Human in the loopLiterally by designMinimal, verifiable tasksRemoved \u2014 runs like softwareNatural outputChat-style assistanceTask execution&#8221;Mainlining&#8221; model intelligence into software<\/p>\n<p>Even the API shape differs, Almeida noted. &#8220;The shape of the API for RLHF is different from RLVR, which is different from what we are doing.&#8221; The team is thinking from scratch, redesigning the stack for reliability rather than engagement. He reminded the audience that genuinely new post-training paradigms tend to &#8220;look totally alien, and then in hindsight become super obvious.&#8221; No one thought about instruction following, he pointed out, until his team made it happen with InstructGPT.<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/08\/58ea28526de2d7e9_1785555987_inline_4.jpg\" alt=\"\"\/><\/p>\n<p>Smarter software, not just cheaper software<\/p>\n<p>Almeida describes himself as a lover of software, and his complaint is that the software industry has not changed in the LLM era. SaaS products have been effectively static since 2019 \u2014 through the entire period of breakneck AI progress \u2014 except where a chatbot has been &#8220;latched on&#8221; to an existing product. That is predictable when AI is assistance-native: the only thing AI enables in SaaS is an assistant on the side.<\/p>\n<p>It is not what early AI pioneers intended. The original OpenAI charter spoke of doing &#8220;tons of work,&#8221; not making profit or shipping chat interfaces. The field expected software to get smarter. Instead, it has gotten cheaper to write \u2014 but the expressibility of the software itself has stayed the same.<\/p>\n<p>Garry Tan&#8217;s phrase &#8220;the golden age of just-in-time software&#8221; is meant as a compliment to the current wave of AI coding tools. Almeida reads it as double-edged. &#8220;I don&#8217;t just want just-in-time software, which is cool,&#8221; he said. &#8220;What I want is smarter software.&#8221; B2B software whose fundamental building blocks have not changed since 2019 is not smarter software. It is the same software written faster.<\/p>\n<p>Automation, in his definition, is the class of rote work so simple it can be communicated once and repeated for basically free \u2014 exactly what computers were supposed to do. Today&#8217;s AI industry is only automating the writing of software. The software itself remains assistance-era.<\/p>\n<p>He predicts the field will eventually write off the current scaling-centric period as &#8220;a weird detour&#8221; \u2014 one the industry did not expect and that largely traces back to those minor algorithmic decisions made years ago.<\/p>\n<p>The scaling law dispute: task matters more than compute<\/p>\n<p>Almeida&#8217;s full-stack philosophy revises Rich Sutton&#8217;s famous &#8220;bitter lesson.&#8221; The bitter lesson holds that algorithms that leverage compute ultimately beat algorithms that leverage human knowledge. That is true in games, Almeida argues, but not in reality. The correct hierarchy, in his view: data matters more than compute, and doing the right task matters more than data.<\/p>\n<p>This is not an abstract philosophical position. It is a direct challenge to the scaling orthodoxy that has priced the current AI boom \u2014 the logic that more compute, extrapolated forward, yields more capability. If task selection and data quality dominate compute in the hierarchy, the capex extrapolation logic that has driven valuations needs revisiting.<\/p>\n<p>Almeida went further in his talk, teasing that the original scaling laws were incorrect. He did not elaborate on the claim, but the implication is clear: someone who co-authored the post-training framework that enabled the scaling era is now arguing that the era was built on miscalibrated assumptions. He promised his newly started Twitter account would post &#8220;very spicy&#8221; things on the subject. For investors and strategists, this is the claim to watch \u2014 a revision to scaling laws from an insider would challenge the compute-first allocation logic that has dominated AI investment for years.<\/p>\n<p>On the Yoshua Bengio suggestion of training a classifier head during pre-training to address hallucination, Almeida is clear: pre-training is not the problem. Pre-trained models are &#8220;incredibly intelligent.&#8221; The problem is entirely in how the field unearths that intelligence. Hallucination comes from the preference-optimization layer, not from the base model. Changing pre-training would not fix it. Changing the post-training objective would.<\/p>\n<p>What TypeSafe is building, and what comes next<\/p>\n<p>TypeSafe remains in stealth but is releasing soon, and the talk doubled as a recruitment pitch: a mailing list for early builders, a careers page, and Almeida&#8217;s public Twitter presence. The company is redesigning the AI stack from scratch for reliability and automation, aiming to move beyond the RLHF assistance paradigm toward calibrated decision-making.<\/p>\n<p>The commercial argument is implied rather than stated outright: value currently concentrated in chat interfaces will migrate to unattended systems that make defensible decisions, and the companies that control post-training objectives \u2014 not just model scale \u2014 will own that migration. Almeida freely concedes that ChatGPT and Claude Code are world-changing products and says he would keep using them. The assistance paradigm will persist. The frontier, he argues, simply moves past it.<\/p>\n<p>For investors watching the space, three signals deserve attention. First, TypeSafe&#8217;s release and whether its approach demonstrates meaningful differences from RLHF and RLVR in real deployment. Second, Almeida&#8217;s promised correction to the scaling laws \u2014 if compute is less predictive of capability than the market has priced, the implications extend well beyond one startup. Third, whether &#8220;calibrated decision-making&#8221; matures into a recognizable post-training paradigm the way RLHF did after InstructGPT, which Almeida himself helped build. The self-described skeptic in the audience got the last word in the room. The rebuttal, Almeida implied, will be empirical \u2014 and it is coming soon.<\/p>\n","protected":false},"excerpt":{"rendered":"Diogo Almeida was one of the architects of RLHF, the training technique that made ChatGPT possible. As a&hellip;\n","protected":false},"author":2,"featured_media":126627,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[7],"tags":[53,580,2798,63426,9396,63428,157,7715,63429,63427],"class_list":["post-126626","post","type-post","status-publish","format-standard","has-post-thumbnail","category-openai","tag-anthropic","tag-chatgpt","tag-claude-code","tag-diogo-almeida","tag-gpt-4","tag-instructgpt","tag-openai","tag-rlhf","tag-rlvr","tag-typesafe-ai"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/126626","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=126626"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/126626\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/126627"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=126626"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=126626"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=126626"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}