{"id":125309,"date":"2026-07-31T04:41:09","date_gmt":"2026-07-31T04:41:09","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/125309\/"},"modified":"2026-07-31T04:41:09","modified_gmt":"2026-07-31T04:41:09","slug":"openai-has-successfully-tripled-its-arc-agi-3-score-by-improving-the-gpt-5-6-harness-demonstrating-that-not-only-the-models-performance-but-also-the-harness-is-important","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/125309\/","title":{"rendered":"OpenAI has successfully tripled its ARC-AGI-3 score by improving the GPT-5.6 harness, demonstrating that not only the model&#8217;s performance but also the harness is important."},"content":{"rendered":"<p>          Jul 31, 2026  12:06:00<\/p>\n<p>         <img decoding=\"async\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/07\/1785472868_22_00_m.png\" border=\"0\" class=\"lzsmall\"\/><a href=\"https:\/\/gigazine.net\/news\/20260326-arc-agi-3\/\" target=\"_blank\" rel=\"nofollow noopener\">ARC-AGI-3<\/a> is a benchmark that measures AI intelligence using games with unknown rules. OpenAI has explained how they tripled their score on this ARC-AGI-3 benchmark.<\/p>\n<p>How enabling two settings tripled our scores on the ARC-AGI-3 benchmark | OpenAI<br \/><a href=\"https:\/\/openai.com\/index\/how-two-settings-tripled-our-arc-agi-3-scores\/\" target=\"_blank\" rel=\"nofollow noopener\">https:\/\/openai.com\/index\/how-two-settings-tripled-our-arc-agi-3-scores\/<\/a><\/p>\n<p lang=\"en\" dir=\"ltr\"> GPT-5.6 Sol has been used to solve open problems in mathematics. So why was it struggling with ARC-AGI-3, a benchmark of 2D puzzle games?<\/p>\n<p>We investigated. The harness was not letting it remember what it had learned.<\/p>\n<p>We found that enabling two API settings tripled our scores\u2026 <a href=\"https:\/\/t.co\/0mv05MnSn2\" rel=\"nofollow\">pic.twitter.com\/0mv05MnSn2<\/a><\/p>\n<p> \u2014 OpenAI (@OpenAI) <a href=\"https:\/\/x.com\/OpenAI\/status\/2082616636989952217?ref_src=twsrc%5Etfw\" rel=\"nofollow\">July 29, 2026<\/a><\/p>\n<p>OpenAI announced the GPT-5.6 series at the end of June 2026. Of these, the highest-performing flagship model is &#8216;GPT-5.6 Sol&#8217;.<\/p>\n<p><a href=\"https:\/\/gigazine.net\/news\/20260629-gpt-5-6-sol-terra-luna\/\" target=\"_blank\" rel=\"nofollow noopener\">OpenAI announces &#8216;GPT-5.6&#8217; series, surpassing Claude Mythos 5, but limited preview release due to US government directive &#8211; GIGAZINE<br \/><\/a><\/p>\n<p><img decoding=\"async\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/07\/1785472869_233_00_m.png\" border=\"0\" alt=\"\" class=\"lzsmall\"\/><\/p>\n<p>GPT-5.6 Sol <a href=\"https:\/\/cdn.openai.com\/pdf\/04d1d1e4-bc75-476a-97cf-49055cd98d31\/cdc_proof.pdf\" target=\"_blank\" rel=\"nofollow noopener\">has solved long-standing unsolved problems in mathematics, such as the cycle double cover conjecture<\/a> , and has <a href=\"https:\/\/www.twitch.tv\/gpt_plays_pokemon\" target=\"_blank\" rel=\"nofollow noopener\">even managed to complete<\/a> &#8216; <a href=\"https:\/\/www.pokemon.co.jp\/ex\/switch-frlg\/\" target=\"_blank\" rel=\"nofollow noopener\">Pok\u00e9mon FireRed<\/a> .&#8217; However, in the ARC-AGI-3 benchmark, which uses a 2D puzzle game, GPT-5.6 Sol reportedly only achieved a score of 7.8%.<\/p>\n<p>The OpenAI development team investigated whether the ARC-AGI-3 2D puzzle game was unusually difficult for GPT-5.6 Sol, or if there was some other reason. Benchmarks evaluate less visible elements such as API settings, harness settings, and prompt display. In the case of ARC-AGI-3, enabling the two API settings used in ChatGPT and Codex (persistence inference and compression) tripled the score on the public task set and reduced the number of output tokens to one-sixth.<\/p>\n<p>OpenAI&#8217;s AI models are trained to think using internal reasoning messages before outputting responses or tool calls. These internal reasoning messages are retained as part of the conversation history, and if the conversation becomes too long, they are summarized and used to continue.<\/p>\n<p>In the official harness, the inference was discarded after each GPT-5.6 Sol operation, and previous actions were also deleted as the context was satisfied, meaning the model had to restart the inference process from the beginning multiple times. Therefore, we have extended the retention period for inference messages.<\/p>\n<p lang=\"en\" dir=\"ltr\"> ARC-AGI-3 tests how well models can learn unfamiliar 2D games without instructions.<\/p>\n<p>The standard harness discarded GPT-5.6 Sol&#8217;s reasoning after each move and dropped earlier actions as the context filled up. The model had to keep starting over. <a href=\"https:\/\/t.co\/cLtHYMNqiO\" rel=\"nofollow\">pic.twitter.com\/cLtHYMNqiO<\/a><\/p>\n<p> \u2014 OpenAI (@OpenAI) <a href=\"https:\/\/x.com\/OpenAI\/status\/2082616638625722669?ref_src=twsrc%5Etfw\" rel=\"nofollow\">July 29, 2026<\/a><\/p>\n<p>Another improvement is <a href=\"https:\/\/developers.openai.com\/api\/docs\/guides\/compaction\" target=\"_blank\" rel=\"nofollow noopener\">replacing rolling cuts with compression<\/a> when the number of input tokens to the model exceeds the maximum number of tokens allowed. Rolling cuts involve &#8216;deleting the oldest tokens first,&#8217; but this can lead to the problem that &#8216;the AI model loses previous observations and actions&#8217; by deleting past tokens. Also, with rolling cuts, the maximum number of tokens is always input even after the token limit is exceeded, which means that &#8216;a large part of the task is performed in a wider context window, resulting in a slight decrease in performance.&#8217; Enabling compression instead of rolling cuts allows GPT-5.6 Sol to better retain what it has learned for each game over longer execution times, enabling it to achieve higher scores with fewer output tokens.<\/p>\n<p>Using the official harness, the ARC-AGI-3 score for GPT-5.6 Sol was 13.3%, while using a harness incorporating two improvements\u2014inference preservation and compression\u2014the ARC-AGI-3 score improved to 38.3%, and the output token count was reduced to one-sixth. For comparison, the average score for human testers on the same task was 48%.<br \/><a href=\"https:\/\/i.gzn.jp\/img\/2026\/07\/31\/how-enabling-two-settings-tripled-arc-agi-3-benchmark\/s01.png\" target=\"_blank\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/07\/s01_m.png\" border=\"0\" class=\"lzsmall\"\/><\/a><\/p>\n<p>OpenAI stated, &#8216;Through these experiments, we hope that people will recognize once again that evaluation does not measure the AI model in isolation, but also simultaneously measures various less visible elements such as API configuration, harness design, and prompt display. This is not the first time we have been surprised by a low score in a public benchmark, only to later discover that the evaluation runner was using a generic harness that omitted inference messages.&#8217;<\/p>\n<p>        <script async src=\"https:\/\/platform.twitter.com\/widgets.js\" charset=\"utf-8\"><\/script><\/p>\n","protected":false},"excerpt":{"rendered":"Jul 31, 2026 12:06:00 ARC-AGI-3 is a benchmark that measures AI intelligence using games with unknown rules. OpenAI&hellip;\n","protected":false},"author":2,"featured_media":125310,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[4],"tags":[6744,8440,3013,6526,4376,8439,8441,717,130,8437,1110,66,4176,136,8438],"class_list":["post-125309","post","type-post","status-publish","format-standard","has-post-thumbnail","category-agi","tag-agi","tag-anime","tag-artificial-general-intelligence","tag-blog","tag-food","tag-game","tag-gigazine","tag-hardware","tag-internet","tag-it","tag-mobile","tag-news","tag-note","tag-software","tag-web-service"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/125309","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=125309"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/125309\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/125310"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=125309"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=125309"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=125309"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}