{"id":51280,"date":"2026-05-26T10:41:12","date_gmt":"2026-05-26T10:41:12","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/51280\/"},"modified":"2026-05-26T10:41:12","modified_gmt":"2026-05-26T10:41:12","slug":"google-deepminds-alphaproof-nexus-solves-erdos-problems-as-ai-math-race-moves-beyond-benchmarks","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/51280\/","title":{"rendered":"Google DeepMind\u2019s AlphaProof Nexus Solves Erd\u0151s Problems as AI Math Race Moves Beyond Benchmarks"},"content":{"rendered":"\n<p>TL;DR<\/p>\n<p>   Main Result: DeepMind says AlphaProof Nexus solved nine open Erd\u0151s problems and proved 44 OEIS conjectures. Proof Method: The system uses Lean to verify each proof step and reportedly solved problems for a few hundred dollars each. Claim Boundary: Hassabis said this is still not AGI, keeping the result framed as a narrower research workflow advance.    <\/p>\n<p>Google DeepMind says its new AlphaProof Nexus math framework has solved nine open Erd\u0151s problems and dozens of other formal conjectures, but CEO Demis Hassabis moved quickly to temper expectations, saying the system is \u201cstill not AGI\u201d even as it points toward a more practical role for AI in verified mathematical research.<\/p>\n<p>AlphaProof Nexus combines a large language model, which proposes possible proof strategies, with Lean, a proof assistant that checks each logical step in software. That matters because mathematics is an area where fluent AI answers can still be wrong: a proof only counts if it survives formal verification. The result is not a general-thinking machine, but a more constrained workflow in which AI-generated ideas can be tested, rejected or certified with far more rigor than an ordinary chatbot response.<\/p>\n<p>In a May 21 arXiv preprint, AlphaProof Nexus was credited with <a href=\"https:\/\/arxiv.org\/abs\/2605.22763v1\" target=\"_blank\" rel=\"nofollow noopener noreferrer\">resolving 9 of 353 open Erd\u0151s problems<\/a>. DeepMind also tied the system to 44 OEIS conjectures, a Hilbert-functions question in algebraic geometry, and an improved convex-optimization bound.<\/p>\n<p>The AlphaProof aper authors argue that formal proof search could matter beyond benchmarks.<\/p>\n<p>\u201cAI-driven formal proof search can serve not only to solve problems but to deepen human understanding.\u201d<\/p>\n<p> Research paper authors <\/p>\n<p>The AlphaProof paper also brings a cost claim of a few hundred dollars per solved problem. That figure applies to checked wins rather than to every abandoned branch in the search process, which keeps the number narrower than a generic model-benchmark claim.<\/p>\n<p> Why the Lean Check Matters <\/p>\n<p>AlphaProof Nexus pairs <a href=\"https:\/\/winbuzzer.com\/tag\/gemini-3-1-pro\/\" target=\"_blank\" rel=\"nofollow noopener\">Gemini 3.1 Pro<\/a> with Lean, and Lean shows that environment through examples of <a href=\"https:\/\/lean-lang.org\/\" target=\"_blank\" rel=\"nofollow noopener noreferrer\">complex pattern matching<\/a> and other machine-checked proof tasks. Software verification changes the trust model because success depends on whether the proof survives formal checking, not on whether the answer merely sounds plausible.<\/p>\n<p>Google DeepMind researchers called the loop the \u201cpower of compiler feedback in grounding LLM reasoning.\u201d Each logical step is checked automatically, giving researchers a more inspectable workflow than a plain text answer from a regular large language model.<\/p>\n<p><a href=\"https:\/\/winbuzzer.com\/wp-content\/uploads\/2026\/05\/AlphaProof-Nexus-agent-design-1.jpg\" target=\"_blank\" rel=\"nofollow noopener\"><img fetchpriority=\"high\" alt=\"AlphaProof Nexus agent design\" width=\"634\" height=\"468\" data-wp-editing=\"1\"  nitro-lazy- nitro-lazy-src=\"https:\/\/cdn-chilj.nitrocdn.com\/gYFaTcLxknXlucWgXPjHDdhAuyobJjHx\/assets\/images\/optimized\/rev-7c44adc\/winbuzzer.com\/wp-content\/uploads\/2026\/05\/AlphaProof-Nexus-agent-design-1.jpg\" class=\"aligncenter size-full wp-image-1952323 nitro-lazy\" decoding=\"async\" nitro-lazy-empty=\"\" id=\"MTgxNTo5MDg=-1\" data-nitro-empty-id=\"MTgxNTo5MDg=-1\" src=\"data:image\/svg+xml;base64,PHN2ZyB2aWV3Qm94PSIwIDAgOTM5IDY5MSIgd2lkdGg9IjkzOSIgaGVpZ2h0PSI2OTEiIHhtbG5zPSJodHRwOi8vd3d3LnczLm9yZy8yMDAwL3N2ZyI+PC9zdmc+\"\/><\/a><\/p>\n<p>LLM unreliability in mathematics research is the specific problem the project tries to narrow: <a href=\"https:\/\/www.indiatoday.in\/technology\/news\/story\/google-ai-solves-56-year-old-math-problems-autonomously-but-deepmind-ceo-says-this-is-still-not-agi-2916518-2026-05-25\" target=\"_blank\" rel=\"nofollow noopener noreferrer\">Mathematical statements known as lemmas<\/a> and to proof paths that can dodge the hardest part of a problem. Checked proofs carry more weight than a fluent draft answer.<\/p>\n<p>Verification also narrows what the cost figure means. Spending counts only when a proof survives checking, not when a model produces something persuasive that later fails formal review.<\/p>\n<p>AlphaProof Nexus also proved 44\/492 OEIS conjectures. DeepMind framed those results as bounded formal successes on defined problem sets, not as proof that one system can now handle all of mathematics or meet an AGI benchmark.<\/p>\n<p> Limits, Published Proofs, and Current Use <\/p>\n<p>Even the basic agent that alternated language-model generation with Lean-based verification solved the same nine Erd\u0151s problems, but it did so at higher computational cost. The contrast leaves efficiency and workflow design as the clearest improvement, while formal verification still acts as a filter for which AI-generated proofs deserve human review.<\/p>\n<p>DeepMind also tied AlphaProof Nexus to active work in math research areas including combinatorics, optimization, graph theory, algebraic geometry, and quantum optics. GitHub hosts <a href=\"https:\/\/github.com\/google-deepmind\/alphaproof-nexus-results\" target=\"_blank\" rel=\"nofollow noopener noreferrer\">Lean and natural-language proofs of AlphaProof on GitHub<\/a>, and the public material gives outside researchers something concrete to inspect instead of a closed benchmark summary.<\/p>\n<p>Released materials include matching natural-language proofs across Erd\u0151s problems, OEIS conjectures, additive combinatorics, algebraic geometry, graph theory, optimization, and quantum optics. GitHub contains successful proofs only, and the repository also points readers to broader attempted-problem sets that did not all end in a final proof.<\/p>\n<p>While such solved-example repositories are easier to audit, they still do not reveal the full cost of every failed search path or every unsolved attempt behind the final successes. That limitation keeps the pricing claim useful yet incomplete.<\/p>\n<p> Prior Math-AI Context and Competitive Environment <\/p>\n<p>DeepMind\u2019s earlier AlphaProof system reached silver-medal-level performance at the International Mathematical Olympiad in the <a href=\"https:\/\/winbuzzer.com\/2024\/07\/26\/deepmind-ai-achieves-silver-in-math-olympiad-xcxwbn\/\" target=\"_blank\" rel=\"nofollow noopener\">2024 Olympiad result<\/a>, which gives the new claim a historical baseline. OpenAI added a newer checkpoint on May 12, with <a href=\"https:\/\/openai.com\/index\/model-disproves-discrete-geometry-conjecture\/\" target=\"_blank\" rel=\"nofollow noopener noreferrer\">disproving a central conjecture in discrete geometry<\/a> tied to an Erd\u0151s-linked problem.<\/p>\n<p>Other proof systems remain active in the same category. Another example is the <a href=\"https:\/\/isabelle.in.tum.de\/\" target=\"_blank\" rel=\"nofollow noopener noreferrer\">Isabelle proof assistant<\/a>, showing that formal verification remains a broader tooling race rather than a one-team field.<\/p>\n","protected":false},"excerpt":{"rendered":"TL;DR Main Result: DeepMind says AlphaProof Nexus solved nine open Erd\u0151s problems and proved 44 OEIS conjectures. Proof&hellip;\n","protected":false},"author":2,"featured_media":25317,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[9],"tags":[2403,20003,1276,30376,2765,2608,30377,29569,8968,111,30378,5044,30379,223,3282,132,7543,1430,8526,5809],"class_list":["post-51280","post","type-post","status-publish","format-standard","has-post-thumbnail","category-google","tag-ai-benchmarks","tag-ai-hallucination","tag-ai-models","tag-ai-reasoning-models","tag-ai-research","tag-alphabet-inc","tag-alphaproof","tag-alphaproof-nexus","tag-artificial-general-intelligence-agi","tag-artificial-intelligence-ai","tag-computational-mathematics","tag-deepmind","tag-erdu0151s-problems","tag-generative-ai","tag-github","tag-google","tag-google-deepmind","tag-google-gemini","tag-large-language-models-llms","tag-mathematics"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/51280","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=51280"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/51280\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/25317"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=51280"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=51280"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=51280"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}