NVIDIA DGX Spark: Everything You Need to Know

The NVIDIA DGX Spark is small enough to fit in your hand, yet it can run AI models with up to 200 billion parameters. With enterprise-grade capability in such a compact package, NVIDIA describes it as the world's smallest AI supercomputer.
But what are its actual specs, and what workloads is it best suited for? Here's everything you need to know about the DGX Spark.
What Is the NVIDIA DGX Spark?
Billed as a personal AI supercomputer for the era of autonomous agents, the DGX Spark's compact footprint is one of its biggest selling points. It measures about 6 x 6 x 2 inches, weighs just 2.6 pounds, and occupies barely more desk space than a gaming console while delivering substantially more AI compute than a traditional desktop workstation.
Featuring the GB10 Grace Blackwell Superchip, Spark combines a 20-core Arm CPU with a Blackwell-architecture GPU. It also comes with a complete AI software stack, allowing developers to prototype, fine-tune, and run large language models locally.
NVIDIA DGX Spark Specifications
Here's a quick summary of key DGX Spark specifications:
| Specification | NVIDIA DGX Spark |
|---|---|
| Superchip | NVIDIA GB10 Grace Blackwell |
| GPU | Blackwell architecture |
| CPU | 20-core Arm (10 Cortex-X925 + 10 Cortex-A725) |
| Memory | 128 GB LPDDR5X coherent unified memory |
| Storage | Up to 4 TB PCIe Gen5 NVMe SSD |
| AI performance | Up to 1 petaFLOP (FP4) |
| Networking | NVIDIA ConnectX-7 |
| Operating system | NVIDIA DGX OS |
| Form factor | 150 x 150 x 50.5 mm |
| Power draw | Approximately 140 W |
| Maximum model size | Up to 200B parameters locally |
| Dual-system support | Up to 405B-parameter models |
In addition to these impressive specifications, Spark's architecture is designed to maximize AI performance within a small desktop system.
How the DGX Spark Punches Above Its Weight
One of the DGX Spark's most notable features is its 128 GB of coherent unified memory. Unlike most desktop systems that divide memory between the GPU and CPU, Spark allows both processors to access the same memory pool.
For AI development, this unified memory can let the DGX Spark run models up to 200 billion parameters locally while avoiding the quantization or model sharding that smaller GPUs often require to fit large models into limited VRAM.
You can also connect two Sparks over high-speed ConnectX networking. This roughly doubles capacity, allowing you to run models of around 405 billion parameters.
There is an important caveat. The DGX Spark's 273 GB/s memory bandwidth is modest compared with a data-center GPU. In practice, Spark has the memory capacity to fit large models that other desktop hardware cannot, but it is not built to deliver the highest possible inference throughput.
In short, the DGX Spark is optimized to make large models accessible on the desktop, not to replace a high-throughput inference server. That trade-off can be a good fit for prototyping and experimentation.
Early Challenges and Ongoing Improvements
At launch, Spark received mixed reviews. Early adopters praised its AI capabilities, but reports of latency, throttling, and reboot issues led some shoppers to hesitate before using it for data-intensive work.
NVIDIA has since addressed some of these concerns through software updates. TensorRT-LLM optimizations and speculative decoding have improved performance, with NVIDIA reporting up to a 2.5x throughput increase across a range of benchmarks compared with launch performance.

Image credit: NVIDIA.
Another concern is its price. Originally announced with a $2,999 starting price, Spark launched in late 2025 at $3,999. In February 2026, NVIDIA raised the price to $4,699, citing global memory supply constraints.
Even at its current price, the DGX Spark remains one of the few desktop systems capable of running such large models locally. Whether it is worth the investment depends on the workload and how often you will use that capability.
What the DGX Spark Is Built For
Spark is not designed to replace every AI workstation or server. For a mini AI supercomputer, it delivers a lot of power, but not enough for every use case.
The DGX Spark is a strong fit for:
- Prototyping new AI applications or fine-tuning models before deploying them at scale.
- Building agentic AI workflows that need ongoing testing and iteration or need to run continuously.
- Developing robotics, computer vision, and video analytics applications with NVIDIA's Isaac and Metropolis frameworks.
It may not be the best choice for:
- High-volume production inference, where the memory-bandwidth ceiling becomes a real constraint.
- Sustained, large-scale training workloads better suited to multi-GPU servers or cloud clusters.
- Gaming or general-purpose compute, where conventional desktop hardware can offer better performance and value.
Final Thoughts
The NVIDIA DGX Spark packs a great deal into a compact form factor. With its GB10 Grace Blackwell Superchip, 128 GB of unified memory, and NVIDIA's AI software stack, it gives developers a capable local environment for experimenting with models that would otherwise be difficult to run on a desktop.
If you regularly prototype applications, fine-tune models, or build autonomous agents, Spark can be a capable workstation for local AI development. But not every workload justifies specialized hardware with a significant upfront cost.
On Vast.ai, you can rent high-performance GPUs on demand and pay only for the compute you use. Access hardware such as RTX 5090s, H200s, B200s, and B300s, then scale capacity up or down as your workloads change.
Get started in minutes. Browse thousands of available GPU instances on Vast.ai and find the compute you need today.


