Nvidia unveils DGX Spark 64GB, a lower-memory version of its Grace Blackwell desktop AI system that will start at $4,999 when partner models arrive on October 23. The new configuration targets developers who want to run capable local models and persistent agents without paying for the original system’s 128GB of unified memory.

 

The hardware keeps the GB10 processor, high-speed networking and Nvidia’s software stack, while introducing Nvidia Sync tools that can join two systems into a 128GB cluster. That makes the product both a cheaper entry point and a test of whether small teams will scale local AI by adding compact nodes instead of buying one larger workstation.

 

The launch has four defining details:

  • $4,999 starting price
  • 64GB of coherent unified memory
  • October 23 partner availability
  • Two-system clustering over 200Gbps networking

 

Related Research

 

Nvidia Unveils DGX Spark 64GB at $4,999

The 64GB configuration will be sold by Acer, ASUS, Dell, Gigabyte, HP and MSI rather than directly by Nvidia. Nvidia describes the product as an OEM-exclusive option, while the existing 128GB configuration remains available for workloads that need more memory in a single box.

 

Reducing memory is the central compromise. Nvidia says one 64GB system can run models with as many as 100 billion parameters, although actual capacity depends on precision, quantization, context length and the memory needed by the runtime. The published limit is therefore a platform ceiling, not a guarantee for every 100-billion-parameter model.

 

The new model retains support for Nvidia Agent Toolkit, CUDA-X libraries and Nemotron models, along with widely used runtimes including Ollama, vLLM and PyTorch with CUDA. Nvidia is positioning the machine as an always-on local agent computer rather than a miniature replacement for a training cluster.

 

That distinction matters for the price comparison. A $4,999 desktop is still specialized equipment, but it can appeal to researchers and small development teams that need local control over code or data and want to avoid sending every inference request to a cloud service.

 

Two DGX Spark Systems Pool 128GB Through Nvidia Sync

Every unit includes a ConnectX-7 network interface. Two 64GB systems can be linked directly with a QSFP cable, allowing Nvidia Sync Cluster Assistant to validate the devices, configure the network and route a workload across both nodes without rebuilding the software environment.

 

The resulting cluster pools 128GB of memory and doubles aggregate memory bandwidth. Nvidia says that expands support to models with as many as 200 billion parameters. In its Qwen 3.8 27B test, two clustered 64GB systems delivered up to 1.7 times the performance of a single system.

 

The benchmark should be read narrowly: it is an Nvidia result for one model and workload, not a universal scaling claim. Distributed inference adds communication overhead, and gains will vary with model architecture, batch size, context and how effectively the runtime divides work between the two machines.

 

Nvidia plans to add Sync Model Launcher later in October. The tool will download and launch Qwen 3.8 27B on one system or a cluster, then configure OpenCode to use the model through a browser. That lowers the setup burden for local coding agents, one of the platform’s clearest target uses.

 

GB10 Hardware Stays the Same Outside Memory

The new configuration uses the same Grace Blackwell GB10 design as the larger-memory DGX Spark. Nvidia lists a 20-core Arm CPU, Blackwell GPU, fifth-generation Tensor Cores and up to one petaFLOP of FP4 AI performance, alongside 273GB per second of memory bandwidth.

 

Connectivity includes the 200Gbps ConnectX-7 interface, 10Gb Ethernet, Wi-Fi 7 and four USB-C ports. Systems can include up to 4TB of NVMe storage, while the compact enclosure measures roughly 150 millimeters square and 50.5 millimeters tall.

 

Nvidia’s product page shows how memory now defines the product ladder: one 64GB system is rated for models up to 100 billion parameters, two 64GB systems reach 128GB and up to 200 billion parameters, and larger combinations of 128GB nodes extend the ceiling further.

 

The approach preserves an upgrade path, but it does not make every two-node purchase economical. Two entry systems cost at least $9,998 before cables, storage choices and taxes. Buyers must weigh that against a 128GB machine’s price, availability and simpler single-node operation.

 

Local AI Moves From Experiments to Persistent Agents

The most important change is not the memory reduction by itself. Nvidia is packaging local inference, clustering and agent software as one supported workflow, aiming to make desktop AI capacity expandable in the same way developers scale servers.

 

That could be useful for coding assistants, document analysis, image generation and research agents that need to remain available throughout the day. Local execution can reduce recurring cloud charges and keep sensitive inputs on premises, although teams still have to manage model licenses, security updates and power use.

 

Blender is preparing a downloadable installer for the platform, while Nvidia’s software supports popular open runtimes and models. Those integrations matter because specialized hardware succeeds only when developers can move existing workloads onto it without adopting an isolated toolchain.

 

Availability begins October 23, making partner pricing and supply the next tests. The $4,999 floor establishes Nvidia’s target, but real-world value will depend on whether OEM systems stay near that level and whether Nvidia Sync produces consistent gains across the models developers actually run.