Editorial illustration for NVIDIA's Groq 3 LPX Enters Full Production for AI Agents
Groq 3 LPX Enters Production for Fast AI Agents
NVIDIA's Groq 3 LPX chip has moved into full production, and the company is pairing it with its Vera Rubin NVL72 rack-scale system to chase a specific problem: getting AI agents to generate tokens fast enough for real-world use. The numbers are concrete. Running Gemma 4 31B, an open-source agentic model, on an Artificial Analysis benchmark, the setup hit 3,400 output tokens per second at 100,000-token context lengths, four times faster than the closest competing platform, according to figures released Tuesday.
The rollout isn't happening in isolation. SpaceXAI says it will use NVIDIA Vera CPUs to power its next generation of agentic AI. CoreWeave has already put Spectrum-X Multiplane into production, a networking setup that links Vera Rubin racks through multiple parallel switches. Nebius has signed on as the first AI cloud provider to adopt Groq 3 LPX.
The timing lines up with Hot Chips, the semiconductor conference running this week in Palo Alto, where inference infrastructure has become the central topic as AI workloads shift from training toward reasoning and multi-agent collaboration, pushing systems to handle far larger context windows than before.
Announced today, the NVIDIA Vera Rubin rack-scale system NVIDIA Groq 3 LPX is in full production. In an Artificial Analysis benchmark running Gemma 4 31B, an open source agentic model, it delivered 3,400 output tokens per second for 100,000-token long-context use cases critical to agentic systems, 4x faster than the nearest alternative platform.
Why this matters
For anyone building agentic systems, the pitch here is speed at scale: 3,400 output tokens per second on a 100,000-token workload, according to an Artificial Analysis benchmark cited by NVIDIA. That's a real number worth watching if it holds up under independent testing. But the framing is oddly mixed.
Groq built its name on custom LPU silicon as a distinct competitor to NVIDIA's GPU stack, so a "Groq 3 LPX" folded into Vera Rubin NVL72 as an NVIDIA product raises questions about what's actually shipping, whose hardware it runs on, and whether this is a partnership, a licensing deal, or marketing shorthand. The model referenced, Gemma 4 31B, also isn't one we can independently verify exists yet. Until NVIDIA or Artificial Analysis publishes fuller benchmark methodology and clarifies the Groq relationship, treat the throughput claims as a vendor's opening number, not an industry baseline.
Worth tracking closely once third-party labs get hands-on access.
Common Questions Answered
What is the token generation speed achieved by NVIDIA's Groq 3 LPX with the Vera Rubin NVL72 system?
According to an Artificial Analysis benchmark, the Groq 3 LPX paired with the Vera Rubin NVL72 rack-scale system achieved 3,400 output tokens per second when running Gemma 4 31B on 100,000-token context lengths. This performance is four times faster than the closest competing platform, making it particularly suited for real-world agentic AI applications.
What specific problem does the Groq 3 LPX and Vera Rubin NVL72 combination address for AI agents?
The primary challenge being addressed is enabling AI agents to generate tokens fast enough for practical, real-world deployment. By achieving 3,400 output tokens per second at long context lengths of 100,000 tokens, the system solves the speed bottleneck that has limited agentic system performance in production environments.
Which open-source agentic model was used to benchmark the Groq 3 LPX performance?
The Artificial Analysis benchmark used Gemma 4 31B, an open-source agentic model, to test the Groq 3 LPX performance. This model was chosen to demonstrate the system's capabilities on realistic agentic workloads with extended context windows.
What is the significance of the 100,000-token context length for agentic systems?
Long-context capabilities of 100,000 tokens are critical to agentic systems because they allow AI agents to maintain and process extensive information during complex tasks. The Groq 3 LPX's ability to maintain 3,400 output tokens per second at this context length enables practical deployment of sophisticated AI agents that require substantial contextual awareness.
Further Reading
- With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents - NVIDIA Blog
- NVIDIA Vera Rubin Opens Agentic AI Frontier - NVIDIA Newsroom
- Inside NVIDIA Groq 3 LPX: The Low-Latency Inference Accelerator for the NVIDIA Vera Rubin Platform - NVIDIA Developer Blog
- NVIDIA Vera Rubin NVL72 - NVIDIA
- NVIDIA Unveils Vera Rubin Platform And Groq 3 Integration to Power Agentic AI Factories - HotHardware