NVIDIA DGX Spark: Everything You Need to Know

July 20, 2026
4 Min Read
By Team Vast

The NVIDIA DGX Spark is small enough to fit in your hand, yet it can run AI models with up to 200 billion parameters. With enterprise-grade capability in such a compact package, NVIDIA describes it as the world's smallest AI supercomputer.

But what are its actual specs, and what workloads is it best suited for? Here's everything you need to know about the DGX Spark.

What Is the NVIDIA DGX Spark?

Billed as a personal AI supercomputer for the era of autonomous agents, the DGX Spark's compact footprint is one of its biggest selling points. It measures about 6 x 6 x 2 inches, weighs just 2.6 pounds, and occupies barely more desk space than a gaming console while delivering substantially more AI compute than a traditional desktop workstation.

Featuring the GB10 Grace Blackwell Superchip, Spark combines a 20-core Arm CPU with a Blackwell-architecture GPU. It also comes with a complete AI software stack, allowing developers to prototype, fine-tune, and run large language models locally.

NVIDIA DGX Spark Specifications

Here's a quick summary of key DGX Spark specifications:

SpecificationNVIDIA DGX Spark
SuperchipNVIDIA GB10 Grace Blackwell
GPUBlackwell architecture
CPU20-core Arm (10 Cortex-X925 + 10 Cortex-A725)
Memory128 GB LPDDR5X coherent unified memory
StorageUp to 4 TB PCIe Gen5 NVMe SSD
AI performanceUp to 1 petaFLOP (FP4)
NetworkingNVIDIA ConnectX-7
Operating systemNVIDIA DGX OS
Form factor150 x 150 x 50.5 mm
Power drawApproximately 140 W
Maximum model sizeUp to 200B parameters locally
Dual-system supportUp to 405B-parameter models

In addition to these impressive specifications, Spark's architecture is designed to maximize AI performance within a small desktop system.

How the DGX Spark Punches Above Its Weight

One of the DGX Spark's most notable features is its 128 GB of coherent unified memory. Unlike most desktop systems that divide memory between the GPU and CPU, Spark allows both processors to access the same memory pool.

For AI development, this unified memory can let the DGX Spark run models up to 200 billion parameters locally while avoiding the quantization or model sharding that smaller GPUs often require to fit large models into limited VRAM.

You can also connect two Sparks over high-speed ConnectX networking. This roughly doubles capacity, allowing you to run models of around 405 billion parameters.

There is an important caveat. The DGX Spark's 273 GB/s memory bandwidth is modest compared with a data-center GPU. In practice, Spark has the memory capacity to fit large models that other desktop hardware cannot, but it is not built to deliver the highest possible inference throughput.

In short, the DGX Spark is optimized to make large models accessible on the desktop, not to replace a high-throughput inference server. That trade-off can be a good fit for prototyping and experimentation.

Early Challenges and Ongoing Improvements

At launch, Spark received mixed reviews. Early adopters praised its AI capabilities, but reports of latency, throttling, and reboot issues led some shoppers to hesitate before using it for data-intensive work.

NVIDIA has since addressed some of these concerns through software updates. TensorRT-LLM optimizations and speculative decoding have improved performance, with NVIDIA reporting up to a 2.5x throughput increase across a range of benchmarks compared with launch performance.

NVIDIA DGX Spark performance improvements from software and model optimizations

Image credit: NVIDIA.

Another concern is its price. Originally announced with a $2,999 starting price, Spark launched in late 2025 at $3,999. In February 2026, NVIDIA raised the price to $4,699, citing global memory supply constraints.

Even at its current price, the DGX Spark remains one of the few desktop systems capable of running such large models locally. Whether it is worth the investment depends on the workload and how often you will use that capability.

What the DGX Spark Is Built For

Spark is not designed to replace every AI workstation or server. For a mini AI supercomputer, it delivers a lot of power, but not enough for every use case.

The DGX Spark is a strong fit for:

  • Prototyping new AI applications or fine-tuning models before deploying them at scale.
  • Building agentic AI workflows that need ongoing testing and iteration or need to run continuously.
  • Developing robotics, computer vision, and video analytics applications with NVIDIA's Isaac and Metropolis frameworks.

It may not be the best choice for:

  • High-volume production inference, where the memory-bandwidth ceiling becomes a real constraint.
  • Sustained, large-scale training workloads better suited to multi-GPU servers or cloud clusters.
  • Gaming or general-purpose compute, where conventional desktop hardware can offer better performance and value.

Final Thoughts

The NVIDIA DGX Spark packs a great deal into a compact form factor. With its GB10 Grace Blackwell Superchip, 128 GB of unified memory, and NVIDIA's AI software stack, it gives developers a capable local environment for experimenting with models that would otherwise be difficult to run on a desktop.

If you regularly prototype applications, fine-tune models, or build autonomous agents, Spark can be a capable workstation for local AI development. But not every workload justifies specialized hardware with a significant upfront cost.

On Vast.ai, you can rent high-performance GPUs on demand and pay only for the compute you use. Access hardware such as RTX 5090s, H200s, B200s, and B300s, then scale capacity up or down as your workloads change.

Get started in minutes. Browse thousands of available GPU instances on Vast.ai and find the compute you need today.