{"id":108092,"date":"2026-07-16T12:12:14","date_gmt":"2026-07-16T12:12:14","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/108092\/"},"modified":"2026-07-16T12:12:14","modified_gmt":"2026-07-16T12:12:14","slug":"ex-openai-cto-muratis-thinking-machines-drops-inkling-a-975b-parameter-model-that-leads-us-labs-but-trails-china","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/108092\/","title":{"rendered":"Ex-OpenAI CTO Murati&#8217;s Thinking Machines drops Inkling, a 975B parameter model that leads US labs but trails China"},"content":{"rendered":"<p>Thinking Machines Lab, the startup founded by former OpenAI CTO Mira Murati, has released Inkling, an open-weights model with 975 billion parameters. It&#8217;s built for efficiency and agent-based tasks, but it still trails the best open-source Chinese models in overall performance.<\/p>\n<p>Thinking Machines Lab has shipped its first production-ready language model. Inkling is a Mixture-of-Experts Transformer with 975 billion total parameters, 41 billion of which are active at any given time. It&#8217;s the first model from the startup founded by Mira Murati, the former OpenAI CTO who played a key role in developing ChatGPT.<\/p>\n<p>Fine-tuning as a business model<\/p>\n<p>Unlike many other open-source AI models, Inkling natively handles text, images, and audio and supports a context window of up to one million tokens. The weights are freely available\u00a0<a href=\"https:\/\/huggingface.co\/thinkingmachines\/inkling\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">on Hugging Face<\/a>. Thinking Machines also offers access through Tinker, its platform for adapting AI models to specific tasks.<\/p>\n<p>The company is positioning Inkling as a flexible base model for customization. &#8220;Inkling is not the strongest overall model available today,&#8221; <a href=\"https:\/\/thinkingmachines.ai\/news\/introducing-inkling\/\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">the announcement<\/a>\u00a0states. Thinking Machines expects the mix of multimodal support, efficient processing, and fine-tuning options to set the model apart.<\/p>\n<p>Thinking Machines says it pre-trained Inkling on 45 trillion tokens of public and synthetic text, images, audio recordings, and videos. The training set also\u00a0<a href=\"https:\/\/thinkingmachines.ai\/training-data-documentation\/\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">includes public data that &#8220;may\u00a0be subject to intellectual property protection.&#8221;<\/a> The company used the Chinese AI model Kimi K2.5, among other methods, to generate synthetic data. Kimi K2.5 <a href=\"https:\/\/the-decoder.com\/cursor-quietly-built-its-new-coding-model-on-top-of-chinese-open-source-kimi-k2-5\/\" rel=\"nofollow noopener\" target=\"_blank\">also served as the basis for Cursor&#8217;s coding model<\/a>. More technical details are available in\u00a0<a href=\"https:\/\/thinkingmachines.ai\/model-card\/inkling\/\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">the model card<\/a>.<\/p>\n<p>Inkling leads U.S. open models but trails China&#8217;s best<\/p>\n<p>According to AI benchmarking <a href=\"https:\/\/x.com\/ArtificialAnlys\/status\/2077466590346444939\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">platform Artificial Analysis<\/a>, Inkling debuts with a score of 41 on the Artificial Analysis Intelligence Index. That makes it the leading open-weights model from a U.S. lab. It ranks three points above the previous leader, Nemotron 3 Ultra at 38, and well ahead of Gemma 4 31B at 29 and gpt-oss-120b at 24.<\/p>\n<p><img fetchpriority=\"high\" decoding=\"async\" class=\"wp-image-37835 size-full\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/07\/inkling_AA_overall_performance-scaled-1.jpg\" alt=\"Inkling deb\u00fctiert auf Platz 41 des Artificial Analysis Intelligence Index und ist damit das f\u00fchrende US-Open-Weights-Modell. | Bild: Artificial Analysis\" width=\"2560\" height=\"2004\"\/>Inkling scores 41 on the Artificial Analysis Intelligence Index, making it the leading U.S. open-weights model. | Image: Artificial Analysis<\/p>\n<p>On GDPval-AA v2, an agent-based benchmark that simulates knowledge-work tasks, Inkling reaches an Elo rating of 1,238. It beats Kimi K2.6 at 1,190 and DeepSeek v4 Flash max at 1,189. Inkling also scores 24 percent on the Tau-3 banking benchmark, ahead of Kimi K2.6 at 21 percent and DeepSeek v4 Flash max at 23 percent.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"wp-image-37836 size-full\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/07\/inkling_AA_GDPval_overall_performance-scaled-1.jpg\" alt=\"\" width=\"2560\" height=\"1135\"\/>Inkling outperforms Kimi K2.6 and DeepSeek v4 Flash max on agent-based knowledge-work tasks. | Image: Artificial Analysis<\/p>\n<p>Inkling performs rather poorly on factual accuracy. Artificial Analysis gives the model a score of just +2 on its AA Omniscience benchmark. That puts it below the leading open-weights models, though still above other U.S. models such as Nemotron 3 Ultra at -1. Inkling&#8217;s accuracy is 40 percent, while its hallucination rate is 63 percent. Those results are likely to limit its use in applications that need highly accurate information.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-58822\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/07\/inkling_AA_GDPval_token_performance-scaled-1.jpg\" alt=\"\" width=\"1930\" height=\"2560\"\/>Inkling scores +2 on AA Omniscience, with 40 percent accuracy and a 63 percent hallucination rate. | Image: Artificial Analysis<\/p>\n<p>With a 64K context window, Inkling costs $1.87 per million input tokens and $4.68 per million output tokens. That&#8217;s slightly more than open-source Chinese models such as GLM-5.2 and DeepSeek v4, which offer similar or better performance on text and code tasks. For context windows up to 256,000 tokens, pricing rises to $3.74 for input, $0.748 for cached input, and $9.36 for output.<\/p>\n<p>But Inkling uses fewer output tokens than comparable open-weights models. According to Artificial Analysis, it averages 25,000 output tokens per Intelligence Index task. GLM-5.2 max uses 43,000, Kimi K2.6 uses about 38,000, and DeepSeek v4 Pro max uses about 37,000 tokens on the same tasks.<\/p>\n<p>Thinking Machines says Inkling offers continuously adjustable &#8220;thinking effort.&#8221; Users can choose their preferred balance between cost and performance, reducing token use while maintaining the same result quality.<\/p>\n<p>Inkling-Small beats the larger model on some benchmarks<\/p>\n<p>Thinking Machines is also <a href=\"https:\/\/thinkingmachines.ai\/news\/introducing-inkling\/#inkling-small\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">previewing Inkling-Small<\/a>, a more compact model with 276 billion total parameters and 12 billion active parameters. The smaller model delivers similar or better results than Inkling on several benchmarks.<\/p>\n<p>Inkling-Small scores 88.3 percent on GPQA Diamond, compared with 87.2 percent for Inkling. On the HLE benchmark with tools, it scores 46.6 percent, slightly ahead of Inkling at 46.0 percent. Thinking Machines credits changes to the pre-training data and training process for the results. The company plans to publish the full weights once testing is complete.<\/p>\n<p>\t\t\t\tAI News Without the Hype \u2013 Curated by Humans<\/p>\n<p>\n\t\t\t\t\tSubscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive &#8220;AI Radar&#8221; frontier report six times a year, full archive access, and access to our comment section.\t\t\t\t<\/p>\n<p>\t\t\t\t<a href=\"https:\/\/the-decoder.com\/subscription\/\" class=\"inline-block text-white bg-(--heise-primary) mt-3 hover:bg-blue-800 focus:ring-4 focus:outline-none focus:ring-blue-300 font-medium rounded-sm w-full sm:w-auto  pl-3 pr-3 py-2.5 text-center newsletter-submit-button hover:no-underline\" rel=\"nofollow noopener\" target=\"_blank\"><br \/>\n\t\t\t\t\tSubscribe now\t\t\t\t<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"Thinking Machines Lab, the startup founded by former OpenAI CTO Mira Murati, has released Inkling, an open-weights model&hellip;\n","protected":false},"author":2,"featured_media":108093,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[7],"tags":[55676,157,51808],"class_list":["post-108092","post","type-post","status-publish","format-standard","has-post-thumbnail","category-openai","tag-inkling","tag-openai","tag-thinking-machines"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/108092","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=108092"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/108092\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/108093"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=108092"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=108092"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=108092"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}