New open models supported by NVIDIA allow developers to build local agents for coding, research and private data processing.

NVIDIA is expanding its local AI ecosystem with a new open-weight model and software tools designed to help developers run, customise and fine-tune AI systems on their own hardware.

The August updates cover agentic AI, coding, multimodal applications, video generation and robotics across RTX PCs, DGX Spark, DGX Station and Jetson devices. NVIDIA is also adding support and optimisations for open-weight models developed by other AI companies.

Among its own releases is Nemotron 3.5 Lightning, a 30-billion-parameter mixture-of-experts model designed for always-on AI agents. NVIDIA says the model can generate tokens up to four times faster than comparable open models and complete tasks 30% faster, although the figures come from the company’s own testing.

Developers can fine-tune Nemotron 3.5 Lightning for specialised workflows. NVIDIA’s open-source NeMo Switchyard library can also direct different stages of an agent workflow to models selected according to factors including accuracy, speed and cost.

Meta’s Muse Glimmer is among the third-party models being optimised for NVIDIA hardware. The 30-billion-parameter open-weight model is designed for coding and agentic workloads and can run on a PC equipped with a single RTX 5090.

Local processing can allow agents to work with files, documents, emails and credentials while reducing the need to send information to cloud-based inference services. However, data protection will still depend on how agents, applications and connected tools are configured.

Other models supported across the ecosystem include NVIDIA’s Cosmos 3 Edge, DeepSeek-V4-Flash, MiniMax-H3, Thinking Machines Lab’s Inkling-Small and LTX-2.5 for video generation.

NVIDIA is also making it easier to combine local computing resources. Its Sync Cluster Assistant can configure two or more DGX Spark systems as a cluster, while a resource monitor scheduled for later in August will track CPU and GPU use across individual devices and clusters.

Why does it matter?

Running capable AI models locally can reduce dependence on cloud services and give developers greater control over data, customisation and computing costs. Open-weight models can also be adapted for specialised tasks without relying entirely on closed model providers.

The expansion nevertheless strengthens NVIDIA’s position across the local AI stack, from chips and computers to model optimisation, routing and deployment software. Greater independence from cloud platforms may therefore create deeper dependence on a particular hardware and software ecosystem.

Would you like to learn more about AI, tech and digital diplomacy? If so, ask our Diplo chatbot!