{"id":69119,"date":"2026-06-10T16:55:07","date_gmt":"2026-06-10T16:55:07","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/69119\/"},"modified":"2026-06-10T16:55:07","modified_gmt":"2026-06-10T16:55:07","slug":"nvidia-accelerates-google-deepminds-diffusiongemma-for-local-ai","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/69119\/","title":{"rendered":"NVIDIA Accelerates Google DeepMind\u2019s DiffusionGemma for Local AI"},"content":{"rendered":"<p>Today, Google DeepMind\u00a0released\u00a0DiffusionGemma\u00a0\u2014\u00a0an experimental open\u00a0model built for exceptionally fast text\u00a0generation.\u00a0NVIDIA has\u00a0optimized\u00a0DiffusionGemma to run even faster across NVIDIA GeForce RTX\u00a0GPUs,\u00a0the\u00a0NVIDIA RTX PRO\u00a0platform\u00a0and NVIDIA DGX Spark\u00a0systems, from local PCs to the cloud.\u00a0<\/p>\n<p>Rather than\u00a0generating\u00a0text one word at a time, DiffusionGemma generates\u00a0multiple words in\u00a0parallel\u00a0to output\u00a0whole blocks of text, opening a new,\u00a0low-latency frontier for the\u00a0kind of\u00a0single-user workloads that developers, researchers and AI enthusiasts run every day.\u00a0<\/p>\n<p>Features of the new model include:\u00a0<\/p>\n<p>Parallel generation:\u00a0DiffusionGemma denoises up\u00a0to 256 tokens\u00a0per step instead of\u00a0predicting one at a time.\u00a0<br \/>\nBuilt on Gemma\u00a04:\u00a0DiffusionGemma is built on Gemma 4, a\u00a026-billion-parameter mixture-of-experts model that activates just 3.8 billion parameters per step, pairing a diffusion head with\u00a0Google\u2019s\u00a0Gemma 4 architecture.\u00a0<br \/>\nUp to 4x faster\u00a0performance: The boost means fast text generation, where single-user generation usually stalls \u2014 on local hardware.\u00a0<br \/>\nOpen and local:\u00a0DiffusionGemma is open\u00a0weights\u00a0under a permissive Apache 2.0 license and runs entirely on RTX and DGX Spark \u2014 no cloud, no per-token cost \u2014 with day-zero support in\u00a0<a target=\"_blank\" href=\"https:\/\/huggingface.co\/nvidia\/diffusiongemma-26B-A4B-it-NVFP4\" rel=\"nofollow noopener\">Hugging Face Transformers<\/a>, vLLM and Unsloth.\u00a0<\/p>\n<p>A Different Way to Generate Text\u00a0<\/p>\n<p>Almost every\u00a0large\u00a0language model\u00a0(LLM)\u00a0in wide use today is autoregressive \u2014\u00a0meaning\u00a0it generates text one token at a time, with each\u00a0new word\u00a0depending on the one before it. That sequential process is what makes interactive AI feel like\u00a0it\u2019s\u00a0typing.\u00a0<\/p>\n<p>DiffusionGemma\u00a0takes a different path. Built on the Gemma 4 26B mixture-of-experts\u00a0architecture,\u00a0it generates text the way diffusion models generate images: by starting from noise and refining a whole block\u00a0of text\u00a0at once. Each step denoises up to 256 tokens in parallel rather than emitting a single token and waiting to compute the next.\u00a0<\/p>\n<p>The result is a model that thinks in blocks instead of sequentially. For latency-sensitive, single-user work \u2014 such as interactive chat, agentic loops or on-device assistants that plan and act \u2014 that parallelism translates into responses fast enough to keep pace with how developers think and iterate.<\/p>\n<p>DiffusionGemma Flies on NVIDIA\u00a0GPUs\u00a0<\/p>\n<p>Generating one token at a time is fundamentally a memory-bound problem \u2014 a traditional LLM spends most of its time waiting on memory bandwidth, not doing math, which\u00a0leaves\u00a0a lot of\u00a0compute\u00a0on the table.\u00a0<\/p>\n<p>Diffusion flips the equation. Pulling a full 256-token block through the transformer in parallel is a compute-bound workload \u2014 exactly what NVIDIA GPUs are built for. NVIDIA Tensor Cores accelerate the dense parallel math, and the CUDA software stack lets the model run efficiently from day one without bespoke tuning.\u00a0In short, the\u00a0model\u2019s\u00a0design plays directly to the\u00a0GPU\u2019\u2018s strengths.\u00a0<\/p>\n<p>DiffusionGemma\u00a0delivers\u00a01,000 tokens\/sec on a single NVIDIA H100 Tensor Core GPU,\u00a0150 tokens\/sec on NVIDIA DGX Spark and\u00a0fastest local inference on\u00a0NVIDIA DGX Station\u00a0\u2014\u00a0roughly 4x\u00a0faster than an equivalent autoregressive model running in the same single-user regime.\u00a0<\/p>\n<p>That advantage holds across NVIDIA\u2019s full lineup, running:\u00a0<\/p>\n<p>Locally on\u00a0the\u00a0NVIDIA\u00a0DGX Spark\u00a0deskside\u00a0personal\u00a0AI\u00a0supercomputer\u00a0\u2014 powered by the\u00a0NVIDIA\u00a0GB10 Grace Blackwell Superchip with 128GB of unified memory \u2014 with the preinstalled NVIDIA AI software stack ready for prototyping, fine-tuning and fully local agent workflows.\u00a0<br \/>\nOn NVIDIA RTX PRO 6000 workstations,\u00a0providing\u00a0developers, researchers and AI professionals\u00a0with\u00a0the\u00a0headroom to run local low-latency generation and agentic loops as part of a professional workflow.\u00a0<br \/>\nOn DGX Station,\u00a0delivering best-in-class,\u00a0high-speed inference at up to\u00a0800 tokens\/sec\u00a0for low-latency text generation and agentic loops\u00a0with 748GB of coherent memory.\u00a0<br \/>\nOn GeForce RTX GPUs, with\u00a0llama.cpp support coming soon.\u00a0<\/p>\n<p>The fastest way to start testing and prototyping\u00a0the model is through\u00a0Hugging Face Transformers, which runs DiffusionGemma on a GeForce RTX 5090 or DGX Spark out of the box. For higher-throughput inference, vLLM\u00a0provides day-zero\u00a0serving support.\u00a0\u00a0<\/p>\n<p>For adapting the model to a specific task or domain, fine-tuning is available through\u00a0Unsloth\u00a0and NVIDIA\u00a0NeMo\u00a0framework, with ready-made DGX Spark playbooks to get a local environment running quickly.\u00a0Check out the\u00a0vLLM\u00a0playbooks for\u00a0<a target=\"_blank\" href=\"https:\/\/build.nvidia.com\/spark\/vllm\" rel=\"nofollow noopener\">DGX Spark<\/a>\u00a0,\u00a0<a target=\"_blank\" href=\"https:\/\/build.nvidia.com\/rtx\/vllm\" rel=\"nofollow noopener\">RTX PRO<\/a>\u00a0and\u00a0<a target=\"_blank\" href=\"https:\/\/build.nvidia.com\/station\/vllm\" rel=\"nofollow noopener\">DGX Station<\/a>.\u00a0<\/p>\n<p>Try Diffusion Gemma on Hugging Face\u00a0or test it for free using NVIDIA-hosted\u00a0application programming interfaces\u00a0at\u00a0<a target=\"_blank\" href=\"https:\/\/build.nvidia.com\/\" rel=\"nofollow noopener\">build.nvidia.com<\/a>.\u00a0<\/p>\n<p>Go deeper on the architecture and local deployment\u00a0by\u00a0reading\u00a0the\u00a0<a target=\"_blank\" href=\"https:\/\/developer.nvidia.com\/blog\/?p=118305\" rel=\"nofollow noopener\">NVIDIA technical blog<\/a> and the <a class=\"Hyperlink TrackedChange TrackChangeHyperlinkInstruction SCXW235303664 BCX0\" href=\"https:\/\/blog.google\/innovation-and-ai\/technology\/developers-tools\/diffusion-gemma-faster-text-generation\/\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Google DeepMind announcement<\/a>.<\/p>\n<p>#ICYMI: The Latest From RTX AI Garage\u00a0<\/p>\n<p>\ud83c\udfac\u00a0NVIDIA researchers released SANA-WM, an open\u00a0source world model that turns a single image and a camera path into a minute-long, 720p video with precise 6-DoF control. At just 2.6 billion parameters, its distilled version generates a full 60-second clip in 34 seconds on a single NVIDIA GeForce RTX 5090\u00a0GPU\u00a0using\u00a0the\u00a0NVFP4\u00a0format\u00a0\u2014 delivering up to 36x higher throughput than comparable open models while running on one GPU. Read\u00a0<a target=\"_blank\" href=\"https:\/\/arxiv.org\/pdf\/2605.15178\" rel=\"nofollow noopener\">the paper.<\/a>\u00a0<\/p>\n<p>\ud83d\udee0\ufe0f\u00a0Building Windows agents just got a full toolset\u00a0\u2014\u00a0<a target=\"_blank\" href=\"https:\/\/developer.nvidia.com\/blog\/build-personal-ai-agents-on-windows-pcs-with-new-tools-from-microsoft-and-nvidia\/\" rel=\"nofollow noopener\">NVIDIA and Microsoft<\/a>\u00a0rolled out turnkey agent sandboxing on native Windows \u2014 Microsoft\u00a0eXecution\u00a0Containers plus the NVIDIA\u00a0OpenShell\u00a0runtime \u2014 alongside up to 2x faster agentic inference and native Windows support for Hermes Agent.\u00a0<\/p>\n<p>\ud83e\udd16DGX Spark goes from unboxing to a running agent in minutes\u00a0\u2014 A streamlined NVIDIA\u00a0NemoClaw\u00a0install gets developers to a working local agent fast, with Qwen3.6-35B running up to 2.6x faster on\u00a0vLLM. And the new cluster assistant in NVIDIA Sync links up to four DGX Spark units into one 512GB pool \u2014 enough for ~400-billion-parameter models.\u00a0<\/p>\n<p>Plug in to RTX Spark on\u00a0<a target=\"_blank\" href=\"https:\/\/www.facebook.com\/NVIDIARTXSpark\/\" rel=\"nofollow noopener\">Facebook<\/a>,\u00a0<a target=\"_blank\" href=\"https:\/\/www.instagram.com\/nvidiartxspark\" rel=\"nofollow noopener\">Instagram<\/a>,\u00a0<a target=\"_blank\" href=\"https:\/\/www.tiktok.com\/@nvidiartxspark\" rel=\"nofollow noopener\">TikTok<\/a>\u00a0and\u00a0<a target=\"_blank\" href=\"https:\/\/x.com\/NVIDIARTXSpark\" rel=\"nofollow\">X<\/a>\u00a0\u2014 and stay informed by subscribing to the\u00a0<a target=\"_blank\" href=\"https:\/\/www.nvidia.com\/en-us\/ai-on-rtx\/?modal=subscribe-ai\" rel=\"nofollow noopener\">RTX Spark newsletter<\/a>.\u00a0<\/p>\n<p>See\u00a0<a target=\"_blank\" href=\"https:\/\/www.nvidia.com\/en-eu\/about-nvidia\/terms-of-service\/\" rel=\"nofollow noopener\">notice<\/a>\u00a0regarding software product information.<\/p>\n<p>\t\t<script async src=\"\/\/www.instagram.com\/embed.js\"><\/script><script async src=\"\/\/www.tiktok.com\/embed.js\"><\/script><\/p>\n","protected":false},"excerpt":{"rendered":"Today, Google DeepMind\u00a0released\u00a0DiffusionGemma\u00a0\u2014\u00a0an experimental open\u00a0model built for exceptionally fast text\u00a0generation.\u00a0NVIDIA has\u00a0optimized\u00a0DiffusionGemma to run even faster across NVIDIA GeForce&hellip;\n","protected":false},"author":2,"featured_media":69120,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[9],"tags":[179,25,5044,38523,132,7543,33422,3072,335,23709,33418],"class_list":["post-69119","post","type-post","status-publish","format-standard","has-post-thumbnail","category-google","tag-agentic-ai","tag-artificial-intelligence","tag-deepmind","tag-dgx-spark","tag-google","tag-google-deepmind","tag-local-ai","tag-nvidia-rtx","tag-open-source","tag-rtx-ai-garage","tag-rtx-spark"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/69119","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=69119"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/69119\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/69120"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=69119"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=69119"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=69119"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}