{"id":148116,"date":"2026-08-22T13:01:08","date_gmt":"2026-08-22T13:01:08","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/148116\/"},"modified":"2026-08-22T13:01:08","modified_gmt":"2026-08-22T13:01:08","slug":"theo-googles-gemini-models-are-just-bad-a-brutal-ranking-of-the-ai-worth-paying-for-biggo-finance","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/148116\/","title":{"rendered":"Theo: Google&#8217;s Gemini Models Are &#8216;Just Bad&#8217; \u2014 A Brutal Ranking of the AI Worth Paying For \u2014 BigGo Finance"},"content":{"rendered":"<p>In the high-stakes game of artificial intelligence, most investors and developers are still obsessed with a single question: which model is the smartest? According to veteran developer Theo, speaking on his podcast Theo &#8211; t3.gg, that is exactly the wrong question to ask in 2026. The real battleground isn&#8217;t intelligence\u2014it&#8217;s value.<\/p>\n<p>Theo\u2019s core argument is that the economics of AI have inverted. Google\u2019s Gemini 3.7 Flash looks cheap at 75 cents per million input tokens, but it burns 73,000 tokens to complete a basic task. OpenAI\u2019s 5.6 Sol costs a premium at $30 per million output tokens, yet its surgical precision means a task requires only 28,000 tokens. &#8220;It is a genius that has to be tamed, whereas 5.6 Sol is a slightly dumber robot that does exactly what you tell it,&#8221; Theo noted, framing the divide that separates the AI elite from a disastrous pack of overpriced underperformers.<\/p>\n<p>The New Elite: Fable 5 and the Cult of Sol<\/p>\n<p>The top of the current hierarchy is defined not by marginal gains, but by a qualitative leap in workflow. Theo places Anthropic\u2019s Fable 5 alone in S+ tier and OpenAI\u2019s 5.6 Sol in S tier, with a deliberately empty A tier separating them from the rest of the industry.<\/p>\n<p>The distinction comes down to autonomy. With mid-tier models like Luna or Kimi K3, a developer must micro-manage: spec out builds, review code before PRs, and verify changes. With Fable or Sol, the workflow is fundamentally different. You provide a vague description, approve a plan, and the model handles implementation, computer-use verification, subagent review, and PR filing autonomously.<\/p>\n<p>Theo\u2019s commitment to this elite tier is measurable. He maintains five Claude subscriptions for Fable 5, draining them via a proxy that bounces between accounts. His usage is &#8220;thousands of dollars a week,&#8221; making API pricing for Fable untenable; only the subscription discount makes it viable. Yet Fable is not his default. &#8220;If I had to pick between Fable and Sol, I would pick Sol. It\u2019s the model I default to&#8230; I would miss Sol more than Fable. But Fable is the best model.&#8221;<\/p>\n<p>The difference is reliability versus brilliance. Fable is &#8220;that genius at the company that nobody wants to work with but nobody wants to fire because they are the smartest person there.&#8221; It overlooks things, takes shortcuts, and struggles inexplicably with iOS development. Sol, by contrast, is less creative but dramatically more token-efficient and reliable enough that Theo completed an entire rewrite of the T3 Code mobile app in SwiftUI using a single thread.<\/p>\n<p>The Token Trap: Why Cheap Models Are Expensive<\/p>\n<p>The central economic insight of the episode is that per-token pricing is now a red herring. A model is only as cheap as the number of tokens it wastes.<\/p>\n<p>The data comparing OpenAI\u2019s Sol to Google\u2019s Gemini is damning. On its highest recommended setting, Sol uses roughly 28,000 tokens per task to achieve near-maximum scores. Cranked to &#8220;Max,&#8221; it uses 60,000 tokens. Google\u2019s Gemini 3.7 Flash, however, consumes 73,000 tokens on its lowest reasoning setting. Push it to High reasoning, and consumption spikes to 107,000 tokens while the quality of the output actually degrades.<\/p>\n<p>ModelReasoning SettingTokens per TaskOutput Quality5.6 Sol (OpenAI)High~28,000Near-max5.6 Sol (OpenAI)Max~60,000MaxGemini 3.7 Flash (Google)Low73,000LowerGemini 3.7 Flash (Google)High107,000Worse than Medium<\/p>\n<p>Theo\u2019s verdict on the search giant is absolute: &#8220;Gemini models make no sense for anything, anything ever.&#8221; The frustration is rooted in history. Gemini 2.0 Flash was once an S-tier model\u2014fast and efficient at $0.10 per million input and $0.40 per million output tokens. That efficiency has evaporated. Gemini 3.5 Flash arrived costing more than Pro in practice, and the newest 3.7 Flash is riding an introductory discount that expires on December 31, 2026, doubling its price just as its inefficiency becomes undeniable. &#8220;It\u2019s a bad model. It\u2019s not like tolerable for some things or better for some things. It is just bad.&#8221;<\/p>\n<p>The Mid-Tier: Deception and Missed Opportunities<\/p>\n<p>The middle of the pack is a graveyard of &#8220;technically interesting but practically flawed&#8221; models. OpenAI\u2019s Terra is the most confusing entry. Priced at $12 per million output tokens (down from $15), it sits logically between Sol and Luna. But because it is less token-efficient than Sol, the per-token savings evaporate in real-world use. &#8220;I have never chosen Terra for anything, and I would be surprised if many people do,&#8221; Theo stated.<\/p>\n<p>Anthropic\u2019s Opus 5 warrants a special kind of scorn. It produces text and code that seems superior on the surface\u2014detailed plans and insights that even Fable occasionally misses. The problem arises when you actually run it. Theo captured the frustration with a vivid analogy: &#8220;I&#8217;ve never had such a weird almost like mimic behavior from a model before, where everything seems good, where it looks like a duck, it smells like a duck, quacks like a duck, and then you cook it and it tastes like shit.&#8221; The code is vague and jargony, mimicking competence without delivering substance. Theo admitted his harsh ranking is partly punitive: &#8220;It tricked me into thinking it was good and I\u2019m mad at it for that.&#8221;<\/p>\n<p>Kimi K3, from Moonshot, is the most impressive open-weight model Theo has tried\u2014the first that could handle end-to-end work with good design taste. But its pricing is a trap. Despite being open weight, its license includes a revenue cap of $10 million; any company above that threshold must strike a deal with Moonshot. This forces all hosting providers to match Moonshot\u2019s MSRP of $15 per million output tokens, preventing the price war that usually follows open-weight releases. Theo won a bet with the CEO of Hugging Face on exactly this point.<\/p>\n<p>The value tier remains dominated by OpenAI\u2019s Luna and DeepSeek\u2019s V4 Flash. Luna, at $1.20 per million output tokens (with a 50% discount on OpenRouter), is Theo\u2019s most-used model by call volume, handling background tasks like title generation and data categorization. V4 Flash is technically impressive and runs on two DGX Spark machines, but it lacks vision support. &#8220;I paste images all the time. That\u2019s important to me,&#8221; Theo explained, downgrading it below Luna.<\/p>\n<p>Vision as a Hard Requirement<\/p>\n<p>The episode identifies vision support as a non-negotiable capability floor. Models are no longer judged solely on text output; developers expect to feed images into their workflow seamlessly. This requirement disqualifies otherwise competent models like V4 Flash and GLM 5.3. The lack of vision in a frontier model is, in Theo\u2019s view, &#8220;embarrassing.&#8221;<\/p>\n<p>The Outlook: Waiting for Astra<\/p>\n<p>The episode closes with a look ahead. The only potential threat to the Fable-Sol duopoly on the horizon is OpenAI\u2019s anticipated &#8220;Astra&#8221; model. Theo\u2019s hope is that Astra can close the gap with Fable on code quality and finally produce visually appealing front ends\u2014a persistent weakness across the current frontier lineup.<\/p>\n<p>The broader implication for the industry is stark. Pricing sheets have become disconnected from economic reality. The AI market is bifurcating into a small group of models that enable autonomous delegation and a much larger group of tools that require constant oversight. For developers, the choice is increasingly binary: pay for the efficiency of Sol or the brilliance of Fable, or burn cash on the false economy of the &#8220;cheap&#8221; alternatives. Until a third option emerges, the gap between the S-tier and everyone else is likely to widen further.<\/p>\n","protected":false},"excerpt":{"rendered":"In the high-stakes game of artificial intelligence, most investors and developers are still obsessed with a single question:&hellip;\n","protected":false},"author":2,"featured_media":148117,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[9],"tags":[72654,53,10769,6010,38107,2408,66063,132,1430,56012,9603,157,46565,68637,72655],"class_list":["post-148116","post","type-post","status-publish","format-standard","has-post-thumbnail","category-google","tag-5-6-sol","tag-anthropic","tag-astra","tag-deepseek","tag-fable-5","tag-gemini","tag-gemini-3-7-flash","tag-google","tag-google-gemini","tag-kimi-k3","tag-moonshot-ai","tag-openai","tag-terra","tag-theo","tag-v4-flash"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/148116","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=148116"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/148116\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/148117"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=148116"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=148116"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=148116"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}