NVIDIA ships Groq 3 LPX to power Vera Rubin’s agentic AI
NVIDIA today said its Groq 3 LPX inference accelerator is now in full production for the Vera Rubin NVL72 rack-scale system.
Source: NVIDIA Newsroom · August 24, 2026 at 8:28 PM · AI-assisted report
Single-source
KUALA LUMPUR, 25 AUGUST 2026 —
NVIDIA Expands Vera Rubin AI Inference Platform with Groq 3 LPX for Agentic Systems
Market Impact
PALO ALTO, Aug 24 — NVIDIA has announced the full production of its Groq 3 LPX inference accelerator, integrated into the Vera Rubin NVL72 platform to enhance fast token generation for agentic AI systems. The move underscores a shift in the AI industry from training-focused infrastructure to inference-driven architectures optimized for reasoning, real-time interaction and multi-agent collaboration.
The Vera Rubin platform, unveiled as part of NVIDIA’s full-stack AI strategy, is engineered to meet the demands of agentic AI—systems that generate tokens at scale, process large context windows and interact with other AI agents. In an Artificial Analysis benchmark using the open-source Gemma 4 31B model, the Groq 3 LPX delivered 3,400 output tokens per second for 100,000-token long-context tasks, four times faster than the nearest competing platform.
This performance is critical as agentic AI workloads increasingly require low-latency, high-throughput inference to support real-time decision-making and tool use.
Industry adoption is accelerating. SpaceXAI will power its next-generation agentic AI systems with NVIDIA Vera CPUs, citing their ability to accelerate CPU-intensive tasks such as orchestration, code execution and simulation. CoreWeave has deployed Spectrum-X Multiplane, a high-bandwidth, flat and lossless AI network architecture that connects Vera Rubin racks using multiple parallel switches.
Nebius, a leading AI cloud provider, is the first to adopt the Groq 3 LPX in production, enabling developers to build interactive agentic applications at scale.
For the Malaysian market, where AI adoption in cloud services and enterprise solutions is growing, the Vera Rubin platform offers potential benefits in latency-sensitive applications such as real-time customer service agents, automated coding assistants and enterprise automation tools. The integration of Groq 3 LPX with Vera Rubin NVL72 could enhance the responsiveness of AI-powered platforms deployed by Malaysian cloud providers or hyperscale data centers, particularly in sectors like fintech, healthcare and smart manufacturing.
However, specific market adoption timelines or local deployments have not been disclosed.
Sector-wise, NVIDIA’s push reflects a broader industry transition toward agentic AI, where inference performance—not just training throughput—drives competitive advantage. The Vera Rubin platform combines NVIDIA Rubin GPUs for large-scale context processing with Groq 3 LPX accelerators optimized for low-latency token generation. At the rack scale, up to 256 LP30 accelerators can be connected via direct chip-to-chip links, forming a unified inference engine designed for deterministic, high-efficiency operation.
Spectrum-X Multiplane further enhances this architecture by providing a flat, multi-plane Ethernet network that avoids the complexity and latency of traditional multi-tier network designs.
Looking ahead, NVIDIA indicates that further optimizations and integrations with Vera Rubin NVL72 are expected, promising new levels of throughput and interactivity. The company is highlighting these advancements at the Hot Chips conference in Palo Alto, emphasizing extreme codesign across compute, networking and software as the foundation of next-generation AI factories. As agentic AI systems grow in complexity and scale, infrastructure that can deliver both speed and efficiency—without trade-offs—will be essential.
For Malaysian enterprises and cloud providers evaluating AI infrastructure, NVIDIA’s integrated platform represents a significant step toward building scalable, responsive and economically viable agentic AI solutions.