{"id":110968,"date":"2026-07-19T10:05:16","date_gmt":"2026-07-19T10:05:16","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/110968\/"},"modified":"2026-07-19T10:05:16","modified_gmt":"2026-07-19T10:05:16","slug":"a-new-ai-benchmark-betting-on-the-world-cup","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/110968\/","title":{"rendered":"A New AI Benchmark: Betting on the World Cup."},"content":{"rendered":"<p>How do you measure whether an <a target=\"_self\" class=\"\" href=\"https:\/\/www.businessinsider.com\/ai-giants-learn-hard-truth-modern-internet-anthropic-openai-google-2026-7\" data-track-click=\"{&quot;element_name&quot;:&quot;body_link&quot;,&quot;event&quot;:&quot;tout_click&quot;,&quot;index&quot;:&quot;bi_value_unassigned&quot;,&quot;product_field&quot;:&quot;bi_value_unassigned&quot;}\" rel=\"nofollow noopener\">AI<\/a> is actually good at predicting the future? The CTO of startup Obside emailed me recently with a fascinating real-world benchmark.<\/p>\n<p>Instead of giving AI models another standardized test, Obside had <a target=\"_self\" class=\"\" href=\"https:\/\/www.businessinsider.com\/chatgpt-ads-big-threat-google-openai-similarweb-search-keywords-conversations-2026-5\" data-track-click=\"{&quot;element_name&quot;:&quot;body_link&quot;,&quot;event&quot;:&quot;tout_click&quot;,&quot;index&quot;:&quot;bi_value_unassigned&quot;,&quot;product_field&quot;:&quot;bi_value_unassigned&quot;}\" rel=\"nofollow noopener\">ChatGPT<\/a>, Gemini, Claude, Grok, Mistral, DeepSeek, and Kimi bet on World Cup matches using live Polymarket odds.<\/p>\n<p>An hour before kickoff, each model goes into agent mode, researches the teams, injuries, and other public information, then decides how much of its virtual $10,000 bankroll to wager.<\/p>\n<p>When I checked after the semifinals on Thursday, French open-source darling Mistral led the field, followed by OpenAI&#8217;s GPT 5.5, and DeepSeek&#8217;s V4. Claude Opus 4.8, meanwhile, sat firmly at the bottom, the only model in the red. Maybe <a target=\"_self\" class=\"\" href=\"https:\/\/www.businessinsider.com\/anthropic-web-bots-crawling-referrals-cloudflare-distillation-2026-7\" data-track-click=\"{&quot;element_name&quot;:&quot;body_link&quot;,&quot;event&quot;:&quot;tout_click&quot;,&quot;index&quot;:&quot;bi_value_unassigned&quot;,&quot;product_field&quot;:&quot;bi_value_unassigned&quot;}\" rel=\"nofollow noopener\">Anthropic<\/a>&#8216;s AI is simply <a target=\"_self\" href=\"https:\/\/www.businessinsider.com\/anthropic-web-bots-crawling-referrals-cloudflare-distillation-2026-7\" data-track-click=\"{&quot;element_name&quot;:&quot;body_link&quot;,&quot;event&quot;:&quot;tout_click&quot;,&quot;index&quot;:&quot;bi_value_unassigned&quot;,&quot;product_field&quot;:&quot;bi_value_unassigned&quot;}\" rel=\"nofollow noopener\">too ethical<\/a> to be a gambler? <\/p>\n<p>Check out the <a target=\"_blank\" href=\"https:\/\/worldcup.obside.com\/\" data-track-click=\"{&quot;click_type&quot;:&quot;other&quot;,&quot;element_name&quot;:&quot;body_link&quot;,&quot;event&quot;:&quot;outbound_click&quot;}\" rel=\" nofollow noopener\">current standings here<\/a>.<\/p>\n<p>This fun exercise measures something many AI benchmarks can&#8217;t: judgment under uncertainty. That&#8217;s a theme I&#8217;ve explored before. Last year, I wrote about ChatGPT entering a <a target=\"_self\" href=\"https:\/\/www.businessinsider.com\/chatgpt-economists-secret-prediction-game-openai-2025-12\" data-track-click=\"{&quot;element_name&quot;:&quot;body_link&quot;,&quot;event&quot;:&quot;tout_click&quot;,&quot;index&quot;:&quot;bi_value_unassigned&quot;,&quot;product_field&quot;:&quot;bi_value_unassigned&quot;}\" rel=\"nofollow noopener\">secret forecasting tournament<\/a> run by economists, where it performed no better than the average human contestant.<\/p>\n<p>Betting on soccer isn&#8217;t the same thing, but it&#8217;s another clever way to test whether AI can turn online information into profitable predictions when nobody yet knows the answer.<\/p>\n<p>Sign up for BI&#8217;s Tech Memo newsletter <a target=\"_self\" href=\"https:\/\/www.businessinsider.com\/subscription\/newsletter\/tech-memo\" data-track-click=\"{&quot;element_name&quot;:&quot;body_link&quot;,&quot;event&quot;:&quot;tout_click&quot;,&quot;index&quot;:&quot;bi_value_unassigned&quot;,&quot;product_field&quot;:&quot;bi_value_unassigned&quot;}\" rel=\"nofollow noopener\">here<\/a>. Reach out to me via email at <a target=\"_blank\" href=\"https:\/\/www.businessinsider.com\/mailto:abarr@businessinsider.com\" data-track-click=\"{&quot;click_type&quot;:&quot;other&quot;,&quot;element_name&quot;:&quot;body_link&quot;,&quot;event&quot;:&quot;outbound_click&quot;}\" rel=\" nofollow noopener\">abarr@businessinsider.com<\/a>.<\/p>\n","protected":false},"excerpt":{"rendered":"How do you measure whether an AI is actually good at predicting the future? The CTO of startup&hellip;\n","protected":false},"author":2,"featured_media":110969,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[2],"tags":[24,25,580,182,6010,2408,11380,6364,9120,527,57086,20299,57087,57085,11133,31637],"class_list":["post-110968","post","type-post","status-publish","format-standard","has-post-thumbnail","category-ai","tag-ai","tag-artificial-intelligence","tag-chatgpt","tag-claude","tag-deepseek","tag-gemini","tag-gpt","tag-grok","tag-mistral","tag-model","tag-new-ai-benchmark","tag-soccer","tag-standardized-test","tag-startup-obside","tag-thursday","tag-world-cup"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/110968","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=110968"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/110968\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/110969"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=110968"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=110968"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=110968"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}