Skip to main content
NVIDIA's sixth-gen NVLink powers millions of AI chips, showcasing advanced data center technology.

Editorial illustration for NVIDIA's Sixth-Gen NVLink Powers Millions of AI Chips

NVIDIA's NVLink Powers AI Factory Data Centers

NVIDIA's Sixth-Gen NVLink Powers Millions of AI Chips

4 min read

NVIDIA has shipped six generations of NVLink since the interconnect debuted in 2016, and the technology now sits at the center of how the company builds what it calls AI factories: data centers designed to run as single, continuous machines rather than racks of separate servers. That shift matters because the workloads have outgrown any one chip's math. Trillion-parameter models, mixture-of-experts architectures, and long-context reasoning all depend on thousands of accelerators trading data with each other in real time, not just crunching numbers in isolation.

Raw FLOPS on a single GPU stopped being the bottleneck years ago. The bottleneck now is how fast those GPUs can talk to each other, how quickly collective operations finish, and how much of that theoretical compute actually turns into usable throughput.

That's the problem NVLink is built to solve. NVIDIA describes it as the scale-up fabric purpose-built for GPU-to-GPU communication at the density and speed AI factories demand, covering training, inference, and other parallel workloads. With millions of accelerators now depending on it, the design choices behind NVLink's sixth generation carry consequences for every operator betting on this infrastructure.

NVIDIA has demonstrated the impact of high-bandwidth, low-latency scale-up networking in large-scale MoEs. For models such as DeepSeek-R1 and Qwen 235B, as well as a simulated 2T parameter LLM, NVLink delivers up to 2.3X the decode throughput compared to leading off-the-shell (OTS) Ethernet.

Why this matters

NVLink's sixth generation isn't a spec bump, it's an admission that the bottleneck in AI infrastructure has moved. FLOPS stopped being the scarce resource a while back; getting data between chips fast enough to keep those FLOPS busy is the new constraint. NVIDIA quietly building this fabric for nearly a decade, with millions of chips already deployed across it, means the company has spent years locking customers into an interconnect standard while everyone else was focused on chip benchmarks.

For developers and founders building on NVIDIA hardware, that's worth noticing: the switches, cables, and software tooling around NVLink represent a moat that's harder to route around than raw compute. Researchers scaling models should read this as confirmation that system-level design, not just accelerator design, now decides what's actually trainable at scale. The open question is how much of this ecosystem lock-in becomes a tax on anyone not building at NVIDIA's scale, and whether rivals like AMD or emerging interconnect standards can offer a real alternative before NVLink's ecosystem advantage becomes insurmountable.

Common Questions Answered

How does NVIDIA's sixth-generation NVLink improve performance for mixture-of-experts models compared to Ethernet?

According to NVIDIA's demonstrations, sixth-generation NVLink delivers up to 2.3X the decode throughput compared to leading off-the-shelf Ethernet solutions for large-scale mixture-of-experts models like DeepSeek-R1 and Qwen 235B. This significant performance improvement is achieved through NVLink's high-bandwidth, low-latency architecture, which enables thousands of accelerators to efficiently exchange data in AI factories.

What is an AI factory according to NVIDIA's definition?

An AI factory is a data center designed to operate as a single, continuous machine rather than as separate racks of individual servers. This architectural approach allows multiple thousands of accelerators to work together seamlessly through NVLink's interconnect technology to handle massive workloads like trillion-parameter models and long-context reasoning tasks.

Why has the bottleneck in AI infrastructure shifted from FLOPS to data transfer between chips?

Modern AI workloads including trillion-parameter models, mixture-of-experts architectures, and long-context reasoning have grown so large that no single chip can handle the mathematical computations alone. This means the limiting factor is no longer raw computational power (FLOPS) but rather the speed and efficiency of moving data between thousands of accelerators, making high-bandwidth interconnects like NVLink critical to infrastructure performance.

How long has NVIDIA been developing NVLink technology before the sixth generation?

NVIDIA has shipped six generations of NVLink since the interconnect technology first debuted in 2016, meaning the company has been refining and developing this technology for nearly a decade. With millions of chips already deployed across NVLink infrastructure, NVIDIA has established a significant installed base locked into this interconnect standard.

LIVE06:16Audit Framework Identifies Rater State Bias in RLHF Preference Data