NVIDIA is extending its Vera Rubin rack-scale system with new LPX and CPX platforms, aiming to accelerate inference for agentic AI systems and reduce token generation costs. The company announced that NVIDIA Groq 3 LPX is now in full production. In benchmarks using the Gemma 4 31B model for long-context use cases critical to agentic AI, the system achieved 3,400 output tokens per second, reportedly four times faster than competing platforms.
This move signals a significant shift in AI infrastructure priorities. As AI moves from training to reasoning and agentic applications, inference performance, throughput, and economics are becoming paramount. Agentic systems are processing larger context windows and generating more tokens, demanding specialized hardware. NVIDIA's approach centers on "extreme codesign," integrating compute, networking, and inference acceleration into a unified system.
Industry adoption is already underway. SpaceXAI plans to use NVIDIA Vera CPUs for its next generation of agentic AI, covering tasks like orchestration and data processing. CoreWeave has deployed NVIDIA Spectrum-X Multiplane, a networking solution designed for high-bandwidth, flat AI networks. Nebius is the first AI cloud provider to adopt NVIDIA Groq 3 LPX, offering developers faster token generation speeds.
NVIDIA Groq 3 LPX is designed to work with the Vera Rubin NVL72 platform, targeting low latency and high throughput for agentic workloads. It accelerates token generation, a key bottleneck in agentic AI's interactive nature. While Vera Rubin GPUs handle context processing, LPX focuses on latency-sensitive decode operations. This pairing aims to eliminate the traditional trade-off between speed and throughput in inference.
NVIDIA sees the evolving infrastructure as a "token factory," capable of delivering performance, throughput, and economic efficiency. The demand for fast inference is driven by agentic AI systems that communicate with multiple AI systems and access vast amounts of data.
