64 GB DGX Spark: memory, more than the petaflop, determines which AI you can run

4 min read
NVIDIA. Official promotional composite of equipment for local AI, published with the 2 October 2026 announcement. Source.

DGX Spark will have an 64 GB configuration starting 23 October, offered by Acer, ASUS, Dell, Gigabyte, HP, and MSI. NVIDIA announced a starting price of 4.999 dollars for these systems, retaining the GB10 Grace Blackwell superchip and DGX OS. To understand the change from the 128 GB version, the main question is how much space the entire workload needs, in addition to the model.

The platform brings together a twenty-core Arm CPU and a Blackwell GPU with unified memory. NVIDIA’s specifications list 273 GB/s of bandwidth and up to one petaflop of FP4 computation. These are different measures: capacity to hold data, speed to move it, and arithmetic capacity at a given precision. None, on its own, is equivalent to “responses per second.”

A calculation that puts 64 GB into perspective

A model with 100.000 million parameters stored at four bits would need, in an idealized calculation, 50.000 million bytes just for those values: 100.000 million × 4 ÷ 8. That is 50 decimal GB before quantization metadata and other components. The example explains why NVIDIA talks about models with up to 100.000 million parameters, but also why that figure does not describe every way of running them.

Inference requires additional space for context, caches, intermediate results, the operating system, and other applications. A model that fits with a short conversation may run into limits as the context grows or it handles multiple requests. Changing the precision or quantization can reduce memory use, with effects that should be evaluated in terms of quality and speed.

Unified memory makes it easier for the CPU and GPU to access the same dataset. It does not create additional capacity: 64 GB is shared among workloads. A useful comparison between configurations should hold the model, precision, input length, output size, and number of concurrent users constant. Without those conditions, a performance test may be measuring a different task.

Two Sparks: more headroom, with a communication cost

NVIDIA lets you connect two 64 GB units via ConnectX-7 and configure them with Sync Cluster Assistant. The company reports up to 1,7 times the performance of one unit in its test with Qwen 3.8 27B. That is a specific result from the manufacturer, not a rule that any program becomes 70% faster.

Distributing a model involves exchanging data between systems. The software must support that distribution, and the benefit depends on how much work is distributed versus how much time communication takes. Two 64 GB memory pools also do not become a single transparent local memory pool of 128 GB for every application.

Software is also part of the configuration

The current Founders Edition notes list DGX OS 7.5.0, driver 580.159.03, and CUDA 13.0.2. NVIDIA cautions that partner systems may receive updates at different times. The software version should accompany any comparison, because memory management and fixed issues can change the result.

For a local AI project, it is worth measuring time to first token, generation speed, memory use, and quality on your own tasks. The DGX Spark use-case guide lays out three workflows. The announcement’s dollar price remains a launch reference: availability, storage, and support in Mexico depend on each commercial configuration.

By evovo Team