NVIDIA Vera Rubin Inference Gains Speed with LPX, CPX Platforms

3 min read
NVIDIA Vera Rubin Inference Gains Speed with LPX, CPX Platforms
NVIDIA Blog
NVIDIA is extending itsVera Rubinrack-scale system with new LPX and CPX platforms, aiming to accelerate inference for agentic AI systems and reduce token generation costs. The company announced that NVIDIA Groq 3 LPX is now in full production. In benchmarks using the Gemma 4 31B model for long-context use cases critical to agentic AI, the system achieved 3,400 output tokens per second, reportedly four times faster than competing platforms. This move signals a significant shift in AI infrastructure priorities. As AI moves from training to reasoning and agentic applications, inference performance, throughput, and economics are becoming paramount. Agentic systems are processing larger context windows and generating more tokens, demanding specialized hardware. NVIDIA's approach centers on "extreme codesign," integrating compute, networking, and inference acceleration into a unified system. Industry adoption is already underway. SpaceXAI plans to use NVIDIA Vera CPUs for its next generation of agentic AI, covering tasks like orchestration and data processing. CoreWeave has deployed NVIDIA Spectrum-X Multiplane, a networking solution designed for high-bandwidth, flat AI networks. Nebius is the first AI cloud provider to adopt NVIDIA Groq 3 LPX, offering developers faster token generation speeds. NVIDIA Groq 3 LPX is designed to work with the Vera Rubin NVL72 platform, targeting low latency and high throughput for agentic workloads. It accelerates token generation, a key bottleneck in agentic AI's interactive nature. While Vera Rubin GPUs handle context processing, LPX focuses on latency-sensitive decode operations. This pairing aims to eliminate the traditional trade-off between speed and throughput in inference. NVIDIA sees the evolving infrastructure as a "token factory," capable of delivering performance, throughput, and economic efficiency. The demand for fast inference is driven by agentic AI systems that communicate with multiple AI systems and access vast amounts of data. The Spectrum-X Multiplane networking architecture is also highlighted. It allows Ethernet to scale to large cluster sizes, reportedly up to 512,000 GPUs, without the latency and cost associated with traditional multi-tier networks. This is achieved by splitting server connections into independent paths or "planes," managed by hardware on NVIDIA's ConnectX SuperNICs. NVIDIA is also introducing Scale-In infrastructure, powered by BlueField-4 processors and DOCA software, to accelerate network services like security and management within agentic AI factories. This aims to make infrastructure services scale alongside AI compute. Furthermore, NVIDIA NVLink Fusion allows custom silicon, or XPUs, to integrate with NVIDIA's AI infrastructure, offering greater flexibility for hyperscalers and AI-native companies. This approach standardizes GPU and XPU systems on a unified architecture, enabling shared rack footprints and infrastructure components. StartupHub.ai data indicates NVIDIA holds a strong position in the semiconductor market with a score of 82/100, distinguishing it from competitors like Broadcom (83/100) and Celestial AI (72/100). The introduction of specialized inference accelerators like Groq 3 LPX alongside the Vera Rubin platform moves NVIDIA further ahead in the critical AI inference segment, a key area for growth and differentiation in the AI hardware market. This focus on codesign across compute and networking is vital as agentic AI demands more sophisticated and efficient infrastructure.
© 2026 StartupHub.ai. All rights reserved. Do not enter, scrape, copy, reproduce, or republish this article in whole or in part. Use as input to AI training, fine-tuning, retrieval-augmented generation, or any machine-learning system is prohibited without written license. Substantially-similar derivative works will be pursued to the fullest extent of applicable copyright, database, and computer-misuse laws. See our terms.