DGX Spark in practice: what you need for local chat, coding, and agents

5 min read
NVIDIA. Official promotional illustration of a studio environment with DGX Spark; it is not a test conducted by evovo. Source.

A local AI computer needs more than a downloaded model. It needs an engine to run it, an interface for working with it, and a way to evaluate its responses. NVIDIA’s official guides for DGX Spark let you separate these components into three use cases: chat, coding assistance, and agents with tools. That separation helps you size memory, disk space, and permissions before combining everything.

The tables consulted in these guides describe DGX Spark with 128 GB. The new 64 GB configuration announced for October should not automatically be assumed to be validated for every recipe. The model, its variant, and the chosen context are part of the configuration, not secondary details.

A local chat has an interface, an engine, and files

The Open WebUI workflow uses a container that integrates the interface with Ollama. NVIDIA estimates about 7 GB for the container image, plus approximately 15 GB for gpt-oss:20b or 25 GB for qwen3.6:latest. These are storage requirements described in the guide; they are not equivalent to total memory usage during a conversation.

The browser displays the interface, while the Spark performs inference. You can access it from its desktop or from another computer through NVIDIA Sync. Downloading models initially requires a connection; running a local model does not in itself involve querying an external API. However, a search connector or remote tool can send information outside, even if the main response is generated locally.

A useful test would be to summarize a low-sensitivity document of your own and check every claim against the original. It is worth measuring both omissions and fabricated details. A fluent response does not prove fidelity: for this task, it matters that the response preserves figures, names, and conditions, and lets you locate the passage that supports each conclusion.

Coding: the model variant changes the calculation

The coding-agent workflow uses Ollama and Qwen3.6:35b-a3b-mtp-q4_K_M, at approximately 23 GB. The same guide puts the q8_0 variant at about 39 GB and bf16 at around 71 GB. This is a concrete example of why the name of a model family alone is not enough to calculate its requirements.

The proposed test connects a terminal agent to that model to complete a small task. Your own evaluation could ask it to fix a reproducible bug in a test repository, run its checks, and review the changes. Success is measured by correct behavior and relevant modifications; the agent finishing its message or generating lots of files does not prove it solved the problem.

An agent adds actions, not just responses

OpenShell introduces policies that restrict files, network connections, and privileges for OpenClaw. The official example is presented as experimental and recommends a clean environment. Its value is in making explicit the difference between a model that proposes an action and a process that has the permissions to carry it out.

For a file-sorting test, for example, you can work on copies in a restricted folder and review the moves before applying them to the originals. If the task needs to access the Internet, that access should be part of the design; if it does not, granting it unnecessarily expands the agent’s scope.

The Spark’s value becomes apparent when you can repeat these tests on your own system and compare models under the same conditions. Record the version, variant, context, response time, and verifiable result. This lets you decide whether a configuration works for your needs more precisely than the model’s size or a prepared demonstration.

By evovo Team