
NVIDIA Groq 3 LPX is now in production as a token-generation accelerator that extends the Vera Rubin platform for agentic AI inference. Designed to work with NVIDIA Vera Rubin NVL72 systems, it targets low-latency output generation, with Nebius and Groq planning early cloud deployments.
How Groq 3 LPX Handles Token Generation
Agentic AI workloads require systems to process large amounts of context and generate output tokens with low latency. Groq 3 LPX targets the second task. NVIDIA uses interactivity to describe the rate at which a system generates tokens for a user.
Token-generation rate determines how quickly an agent can inspect files, write and test code, call tools, verify results and repeat operations. Vera Rubin NVL72 supports AI training and inference, while Groq 3 LPX specializes in token generation. NVIDIA did not provide token-rate, latency or power-efficiency benchmarks in the announcement.
“Inference is the growth engine of AI. NVIDIA Grace Blackwell and NVL72 revolutionized large language model inference with an unprecedented leap in performance and efficiency,” said Jensen Huang, founder and CEO of NVIDIA. “Vera Rubin extends that vision with workload-optimized AI factory configurations designed for the era of agentic AI, advancing the performance frontier with LPX for ultrafast token generation. This transforms how intelligence is produced, delivering another giant leap in AI throughput, efficiency and responsiveness, just as demand for AI computation is accelerating worldwide.”
Nebius and Groq Plan Cloud Deployments
Nebius plans to add NVIDIA Groq 3 LPX to Nebius Token Factory, its production inference platform. The deployment will serve developers running high-volume inference workloads with strict latency requirements.
“Generation is the phase of inference that determines how responsive an AI system actually is, and that’s exactly what NVIDIA Groq 3 LPX is built to accelerate,” said Danila Shtan, chief technology officer of Nebius. “As the first AI cloud bringing it to production via Nebius Token Factory, we’re making sure every step of an agent’s loop feels instant – through the same API developers are already using, with no migration to a new stack.”
Groq, an AI inference cloud provider, also plans to become an early adopter of Groq 3 LPX and Vera Rubin NVL72.
Vera Rubin Rack Architecture
The Vera Rubin platform spans seven chips and five rack designs. Vera Rubin NVL72 and Groq 3 LPX address the training and inference requirements of frontier model developers and open-model service providers.
The rack systems include NVIDIA BlueField-4 data processing units (DPU). They operate with NVIDIA Vera CPU racks, NVIDIA Vera BlueField-4 STX storage and NVIDIA Spectrum-6 SPX Ethernet to support multi-agent workloads across computing, networking and storage.
Source: NVIDIA
About NVIDIA
![]()
NVIDIA, founded in 1993 and headquartered in Santa Clara, CA, designs and manufactures graphics processing units, systems on chips, networking hardware, and AI intelligence software such as CUDA. Its products serve industries including gaming, data centers, autonomous vehicles, professional visualization, robotics, health care, and energy. The company introduced the GPU in 1999 and later expanded into accelerated computing and AI infrastructure. In gaming, its GPUs support high-performance rendering, while in AI and high-performance computing, its systems provide the infrastructure for training and deploying large-scale models. NVIDIA also develops tools for robotics and autonomous driving.
About Nebius

Nebius is a technology company that provides cloud infrastructure and computing platforms for AI development and deployment. The company offers GPU-based cloud computing, data processing tools, and infrastructure for training and running machine learning models. Nebius serves developers, startups, enterprises, and research organizations that build AI and data-intensive applications. Its platforms support industries such as healthcare, robotics, financial services, retail, and media. The company also operates related businesses including Avride for autonomous mobility and TripleTen for technology education and holds stakes in companies such as Toloka and ClickHouse. Nebius Group N.V. traces its origins to 1989 and later emerged from the restructuring of Yandex’s international operations. The company is headquartered in Amsterdam, Netherlands.
