{"id":94752,"date":"2026-07-03T23:43:20","date_gmt":"2026-07-03T23:43:20","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/94752\/"},"modified":"2026-07-03T23:43:20","modified_gmt":"2026-07-03T23:43:20","slug":"gpt-and-claude-failed-bridgewaters-finance-tests-because-the-right-answers-were-never-public","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/94752\/","title":{"rendered":"GPT and Claude failed Bridgewater&#8217;s finance tests because the right answers were never public"},"content":{"rendered":"<p>Hedge fund Bridgewater and Thinking Machines Lab say a fine-tuned open-weight model outperforms the strongest AI models at evaluating financial documents, at a fraction of the cost. The numbers come from their own internal evaluation.<\/p>\n<p>Investors get buried in news, analysis, corporate filings, and emails every day. According to a report from Bridgewater&#8217;s\u00a0<a href=\"https:\/\/www.bridgewater.com\/aia-labs\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">AIA Labs<\/a>\u00a0and\u00a0<a href=\"https:\/\/thinkingmachines.ai\/news\/learning-to-replicate-expert-judgment-in-financial-tasks\/\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">Thinking Machines Lab<\/a>, the startup founded by former OpenAI CTO Mira Murati, reading isn&#8217;t the real work. The real work is the constant stream of small, repeated judgment calls about what actually matters. That&#8217;s the triage the researchers wanted to automate.<\/p>\n<p>They defined six tasks drawn from an investor&#8217;s daily routine. One example: deciding whether a financial article is relevant to an executive. Another: whether a central bank document signals the direction of future rate changes. For investors, these calls are trivial, but they can barely put their reasoning into words. The report gives a telling example. A headline about Trump&#8217;s claim to Greenland gets flagged as irrelevant, while Trump&#8217;s threat of new China tariffs is highly relevant. Both touch on geopolitics and finance.<\/p>\n<p>Frontier models failed in the authors&#8217; tests. Variants of Gemini, Claude, and GPT hit only about 50 percent accuracy with a basic prompt. Expert-written instructions and a three-tier rating system (&#8220;relevant and interesting,&#8221; &#8220;relevant but uninteresting,&#8221; &#8220;irrelevant&#8221;) pushed accuracy into the mid-70s. That still fell short of the 80 percent threshold the authors set for trustworthy deployment.<\/p>\n<p><img fetchpriority=\"high\" decoding=\"async\" class=\"size-full wp-image-58142\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/07\/Finance-Prompt-native-ThinkingMachines-Bench.png\" alt=\"\" width=\"1130\" height=\"569\"\/>When experts write the prompt, performance jumps sharply compared to a naive prompt. | Image: Thinking Machines<\/p>\n<p>Newer models barely improve per dollar, the report says. GPT 5.4 costs 43 percent more than 5.2 but is only marginally more accurate.<\/p>\n<p>The real value lives inside investors&#8217; heads<\/p>\n<p>The solution was fine-tuning, retraining an open-weight model on proprietary examples. The key ingredient was the Bridgewater investors&#8217; judgment: At first, cheap outside contractors labeled the documents, but many of those labels were wrong. To avoid having expensive professionals review everything, the researchers used a workaround. A first model learned from the flawed labels and re-evaluated the same documents. Wherever the model and the original label disagreed, there was likely an error. Only those disputed cases went to investors for correction.<\/p>\n<p>Training ran on the\u00a0<a href=\"https:\/\/the-decoder.com\/ex-openai-cto-mira-murati-introduces-tinker-an-api-for-fine-tuning-of-open-weight-llms\/\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">Tinker platform<\/a>\u00a0from Thinking Machines Lab, built on top of the open model Qwen3-235B. In the team&#8217;s own evaluation, the fine-tuned model hit 84.7 percent accuracy versus 78.2 percent for the best frontier model tested. It also cost nearly 14 times less to run. This isn&#8217;t a truly independent comparison, of course. Both companies have a clear interest in selling their product.<\/p>\n<p>Still, the finding beyond the numbers is worth noting. It shows once again that the big labs like OpenAI haven&#8217;t absorbed all the data out there. Huge pools of proprietary corporate data and untrained human expertise still exist, and they hold real room for improvement. That&#8217;s especially true where companies deliberately keep their most valuable data private. Anyone who hands that data to a frontier lab risks competing against a product built on top of it.<\/p>\n<p>Fine-tuning open models through tools like Tinker gives companies an alternative. They keep the weights, the data, and, depending on the setup, the GPUs themselves.<\/p>\n<p>\t\t\t\tAI News Without the Hype \u2013 Curated by Humans<\/p>\n<p>\n\t\t\t\t\tSubscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive &#8220;AI Radar&#8221; frontier report six times a year, full archive access, and access to our comment section.\t\t\t\t<\/p>\n<p>\t\t\t\t<a href=\"https:\/\/the-decoder.com\/subscription\/\" class=\"inline-block text-white bg-(--heise-primary) mt-3 hover:bg-blue-800 focus:ring-4 focus:outline-none focus:ring-blue-300 font-medium rounded-sm w-full sm:w-auto  pl-3 pr-3 py-2.5 text-center newsletter-submit-button hover:no-underline\" rel=\"nofollow noopener\" target=\"_blank\"><br \/>\n\t\t\t\t\tSubscribe now\t\t\t\t<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"Hedge fund Bridgewater and Thinking Machines Lab say a fine-tuned open-weight model outperforms the strongest AI models at&hellip;\n","protected":false},"author":2,"featured_media":94753,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[8],"tags":[53,3154,182,1151,8554,49556],"class_list":["post-94752","post","type-post","status-publish","format-standard","has-post-thumbnail","category-anthropic","tag-anthropic","tag-anthropic-claude","tag-claude","tag-finance","tag-thinking-machines-lab","tag-tinker"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/94752","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=94752"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/94752\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/94753"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=94752"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=94752"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=94752"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}