Skip to main content
NVIDIA DGX Spark system with 64GB HBM3, running 100B-parameter AI models on-device, showcasing advanced GPU technology.

Editorial illustration for NVIDIA's New DGX Spark 64GB Runs 100B-Parameter Models On-Device

NVIDIA DGX Spark 64GB Runs 100B Models On-Device

• 4 min read

NVIDIA is shipping a smaller, cheaper version of its DGX Spark desktop AI computer this month, and the number that matters is 64. That's the gigabytes of unified memory in the new configuration, down from the 128GB model that launched earlier, and it still runs models with up to 100 billion parameters entirely on the machine sitting on someone's desk. No cloud, no API bill, no sending a company's data to a third-party server to get an answer.

The 64GB DGX Spark will come from six manufacturer partners, Acer, ASUS, Dell, Gigabyte, HP and MSI, each preloading NVIDIA's DGX OS and full AI software stack so the machine works the moment it's unboxed. Under the hood it keeps the same GB10 Grace Blackwell Superchip and ConnectX-7 networking found in the pricier 128GB version, just with less memory to hit a lower price.

That tradeoff raises an obvious question for developers: what happens when a single 64GB box isn't enough for a bigger model or a heavier agentic workload. NVIDIA built an answer into the hardware itself.

The new 64GB configuration, available exclusively from manufacturer partners, keeps the platform at an accessible price point while retaining the GB10 Grace Blackwell Superchip, DGX OS and full NVIDIA AI software stack — same as the 128GB model. It supports up to 100-billion-parameter models and the agentic applications built on them, fully on device.

Why this matters

The 64GB DGX Spark is NVIDIA betting that local inference, not cloud rental, becomes the default way developers prototype with large models. Running 100-billion-parameter models on a desk unit from Acer, ASUS, Dell, Gigabyte, HP or MSI changes the calculus for anyone who's been budgeting GPU-hours just to test an agent pipeline. For founders, that's a real cost argument: build and iterate on-device, then decide if you actually need cloud scale for production.

We'd push back on one thing: "ready to use" with DGX OS and NVIDIA's stack still means you're locked into NVIDIA's software world, which has its own tradeoffs for teams wanting portability. And multi-device setups for bigger models mean this isn't a one-box solution for everyone, despite the framing.

Still, the timing is notable. Open models keep shrinking to fit consumer-adjacent hardware, and NVIDIA is clearly positioning Spark as the on-ramp for that trend. Worth watching: how quickly third-party benchmarks confirm real-world performance once these ship this month, rather than NVIDIA's own use-case framing.

Common Questions Answered

What is the key difference between the new 64GB DGX Spark and the original 128GB model?

The new 64GB DGX Spark configuration reduces unified memory from 128GB to 64GB, making it smaller and more affordable while maintaining the same GB10 Grace Blackwell Superchip and full NVIDIA AI software stack. Despite the reduced memory, it still supports running up to 100-billion-parameter models entirely on-device without requiring cloud infrastructure.

Can the 64GB DGX Spark run 100-billion-parameter models without cloud connectivity?

Yes, the 64GB DGX Spark can run models with up to 100 billion parameters completely on-device without any cloud connection or API bills. This allows users to keep their company data local and avoid sending it to third-party servers while testing and prototyping large language models.

Which manufacturers are offering the 64GB DGX Spark configuration?

The 64GB DGX Spark is available exclusively from six manufacturer partners: Acer, ASUS, Dell, Gigabyte, HP, and MSI. This multi-vendor approach makes the platform more accessible to developers and organizations looking for on-device AI inference capabilities.

How does local inference on the DGX Spark change the development process for AI applications?

Running 100-billion-parameter models locally on the DGX Spark eliminates the need to budget for GPU-hours in the cloud just to prototype and test agent pipelines. Developers can now build and iterate on-device at their own pace, then decide whether they actually need cloud-scale infrastructure for production deployment, reducing both costs and iteration time.

LIVE17:22Gemini 3.8 AI Models Promise Better Reasoning at Lower Cost