
SANTA CLARA, CA (AI Infra Summit), Sep 11, 2025 – NVIDIA has announced the NVIDIA Rubin CPX, a GPU designed for large-context processing, intended to support AI systems manage million-token code generation and generative video tasks with improved computational efficiency.
Rubin CPX works with NVIDIA Vera CPUs and Rubin GPUs inside the NVIDIA Vera Rubin NVL144 CPX platform. The NVIDIA MGX system packs 8 exaflops of AI compute to provide 7.5x more AI performance than NVIDIA GB300 NVL72 systems, as well as 100TB of memory and 1.7 petabytes per second of memory bandwidth in a single rack. A Rubin CPX compute tray will also be offered for customers looking to reuse existing Vera Rubin NVL144 systems.
“The Vera Rubin platform will mark another leap in the frontier of AI computing – introducing both the next-generation Rubin GPU and a new category of processors called CPX,” said Jensen Huang, founder and CEO of NVIDIA. “Just as RTX revolutionized graphics and physical AI, Rubin CPX is the first CUDA GPU purpose-built for massive-context AI, where models reason across millions of tokens of knowledge at once.”
NVIDIA Rubin CPX delivers high performance and efficient token handling for long-context workloads, exceeding the limits of current systems. The platform is designed to move AI coding assistants beyond simple code generation to tools that can optimize large software projects.
Processing an hour of video can require up to 1 million tokens, which strains conventional GPU setups. Rubin CPX combines video decoders and encoders with long-context inference processing on a single chip to handle these workloads, enabling applications like video search and high-quality generative video.
Built on the NVIDIA Rubin architecture, the Rubin CPX GPU uses a monolithic die design. It includes NVFP4 computing resources and is optimized to deliver high performance and energy efficiency for AI inference tasks.
Advancements Offered by Rubin CPX
Rubin CPX delivers up to 30 petaflops of compute with NVFP4 precision for the highest performance and accuracy. It features 128GB GDDR7 memory to accelerate context-based workloads. In addition, it delivers 3x faster attention capabilities compared with NVIDIA GB300 NVL72 systems.
Rubin CPX is offered in multiple configurations, including the Vera Rubin NVL144 CPX, that can be combined with the NVIDIA Quantum‑X800 InfiniBand scale-out compute fabric or the NVIDIA Spectrum-X Ethernet networking platform with NVIDIA Spectrum-XGS Ethernet technology and NVIDIA ConnectX-9 SuperNICs. Vera Rubin NVL144 CPX enables companies to monetize, with $5 billion in token revenue for every $100 million invested.
Industry Leaders Look to Rubin CPX
AI innovators are exploring how Rubin CPX can accelerate their applications, ranging from large-scale software development to the analysis of visual content to understand moving images.
Cursor sees the benefits of Rubin CPX to boost developer productivity with code generation and collaborative tools in the coding environment. “With NVIDIA Rubin CPX, Cursor will be able to deliver lightning-fast code generation and developer insights, transforming software creation,” said Michael Truell, CEO of Cursor. “This will unlock new levels of productivity and empower users to ship ideas once out of reach.”
Runway will use NVIDIA technologies to enable creators to produce cinematic content and visual effects. “Video generation is rapidly advancing toward longer context and more flexible, agent-driven creative workflows,” said Cristóbal Valenzuela, CEO of Runway. “We see Rubin CPX as a major leap in performance, supporting these demanding workloads to build more general, intelligent creative tools. This means creators – from independent artists to major studios – can gain unprecedented speed, realism and control in their work.”
Magic is an AI research and product company developing foundation models to power AI agents that can automate software engineering. “With a 100-million-token context window, our models can see a codebase, years of interaction history, documentation and libraries in context without fine-tuning,” said Eric Steinberger, CEO of Magic. “This enables users to coach the agent at test time through conversation and access to their environments, bringing us closer to autonomous agentic experiences. Using a GPU like NVIDIA Rubin CPX greatly accelerates our compute workloads.”
Software Support
NVIDIA Rubin CPX will be supported by the complete NVIDIA AI stack – from accelerated infrastructure to enterprise‑ready software. The NVIDIA Dynamo platform scales AI inference, boosting throughput while reducing response times and model serving costs.
The processors will be able to run the latest in the NVIDIA Nemotron family of multimodal models that provide reasoning for enterprise-ready AI agents. For production-grade AI, Nemotron models can be delivered with NVIDIA AI Enterprise.
The Rubin platform extends NVIDIA’s developer ecosystem – with NVIDIA CUDA‑X libraries, a community of over 6 million developers and nearly 6,000 CUDA applications.
Availability
NVIDIA Rubin CPX is expected to be available at the end of 2026.
Source: NVIDIA
About NVIDIA
![]()
NVIDIA Corporation, founded in 1993 and based in Santa Clara, CA, designs and produces graphics processing units, systems on chips, networking hardware and AI software such as CUDA. Its hardware and software support applications in gaming, data centers, autonomous vehicles, professional visualization, robotics, health care and energy. The company pioneered the GPU in 1999 and later expanded into accelerated computing and AI infrastructure. In gaming, its GPUs drive high-performance rendering, while in AI and high-performance computing, its systems provide the infrastructure for training and deploying large-scale models. Nvidia also develops tools for robotics and autonomous driving systems. For the fiscal quarter ending January 2025, it reported revenue of $39.3 billion and net income of $22.1 billion.