Featured image of post NVIDIA Unveils Full-stack Upgrade for Vera Rubin Ecosystem: Groq 3 LPX, Spectrum-X Multi-Plane, and Scale-In Define New AI Factory Architecture

NVIDIA Unveils Full-stack Upgrade for Vera Rubin Ecosystem: Groq 3 LPX, Spectrum-X Multi-Plane, and Scale-In Define New AI Factory Architecture

At Hot Chips, NVIDIA reveals five core technologies for Vera Rubin ecosystem, defining a complete AI factory with end-to-end capabilities.

NVIDIA Unveils Five Core Technologies for Vera Rubin AI Factory Ecosystem

NVIDIA Unveils Five Core Technologies for Vera Rubin AI Factory Ecosystem
NVIDIA Unveils Five Core Technologies for Vera Rubin AI Factory Ecosystem|News screenshot

On August 24, 2024 (local time), at the Hot Chips conference, NVIDIA announced major advancements in the Vera Rubin AI factory ecosystem:

  • NVIDIA Groq 3 LPX: Entered full-volume production as a low-latency inference accelerator for token generation
  • Spectrum-X Multi-Plane: Enables scaling to 512,000 GPUs without adding Layer-3 networking
  • Scale-In: Based on BlueField-4 and DOCA, handles security, storage access, and operations offload
  • NVLink Fusion: Enables third-party XPU and CPU integration into NVIDIA’s rack architecture

Initial Deployment: Nebius becomes the first adopter, integrating Groq 3 LPX into its production inference platform Nebius Token Factory, allowing developers to retain their existing API stack without migration.

Groq 3 LPX: Low-Latency Token Generation Accelerator

Groq 3 LPX is explicitly defined by NVIDIA as an “interactive AI inference accelerator”—not a replacement for Vera Rubin GPU but as an extension that specifically accelerates token generation phase.

Applications like code agents involve hundreds to thousands of execution steps, where per-generation latency accumulates, directly impacting end-to-end task completion time. Test results show:

  • Output rate of 3,400 tokens/second on Gemma 4 31B with 100K token context in Artificial Analysis benchmark, the fastest recorded for this model
  • Response speed up to 4x faster than closest alternative platforms on agent and latency-sensitive workloads

Deployment will prioritize AI cloud providers, with Nebius first and Groq also planned as an early adopter.

Spectrum-X Multi-Plane: Scaling to 512K GPUs Without Layer-3 Overhead

Spectrum-X Multi-Plane: Scaling to 512K GPUs Without Layer-3 Overhead
Spectrum-X Multi-Plane: Scaling to 512K GPUs Without Layer-3 Overhead|News screenshot

Legacy two-layer Ethernet clusters often require adding Layer-3 networking whenscaling beyond certain sizes, incurring additional hops, latency, jitter, and hardware costs.

The surprising achievement of Spectrum-X multi-plane: In an eight-plane topology, one plane failure still retains ~90% total bandwidth, with hardware-based recovery 11x faster than software alternatives, resulting in 1.6x AI factory output improvement.

Key technical specifications:

  • Spectrum-X SN6000 switch uses 102.4Tb/s Spectrum-6 ASIC
  • ConnectX-9 SuperNIC delivers up to 1,600Gb/s bandwidth per GPU
  • Hardware engine in SuperNIC handles traffic allocation and failover
  • Spectrum-XGS connects multiple data centers, improving multi-site NCCL collective performance by 1.9x

Scale-In is dubbed the “fifth pillar” of NVIDIA’s AI networking architecture. Powered by BlueField-4 processor and DOCA software stack, it isolates multi-tenant networking, storage access, security, resource provisioning, and real-time observability into a dedicated hardware domain, preventing continuous CPU/GPU resource consumption.

Complementing this is NVLink Fusion, which extends NVLink beyond NVIDIA GPU interconnect to third-party XPU and CPU via sixth-generation NVLink, NVLink Switch, and NVLink-C2C:

  • Sixth-generation NVLink supports 72-XPU interconnect domains with endpoint-to-endpoint latency reduced to one-third and packet rate increased 10x versus general Ethernet solutions
  • NVLink-C2C connects XPU to NVIDIA Vera CPU or ecosystem CPUs with up to 6x power efficiency over PCIe

This enables hyperscalers to deploy custom XPU alongside NVIDIA GPUs in existing MGX rack infrastructure (power, cooling, networking, management software, supply chain), allowing dynamic chip mix adjustments based on supply and workload.

Key Technical Comparison

Key Technical Comparison
Key Technical Comparison|News screenshot

TechnologyCore FunctionKey MetricBaselineImprovement
Groq 3 LPXToken generation acceleration3,400 token/s (Gemma 4 31B)Closest alternative4x faster response
Spectrum-X multi-planeLarge-scale cluster scaling512K GPUs without Layer-3Traditional 3-layer11x faster recovery, 1.6x output
NVLink FusionXPU integration72-XPU interconnect domainEthernet-based1/3 latency, 10x packet rate
Scale-InInfrastructure service offloadBlueField-4 + DOCAHost CPU/GPUIsolates security/storage/Ops

##落地建议 (Application Guidance)

Suitable for immediate adoption: Cloud providers or enterprise AI platforms with existing large-scale GPU clusters facing token generation latency or network scaling bottlenecks should evaluate the Groq 3 LPX plus Spectrum-X multi-plane combination. Hyperscalers already using MGX racks and planning custom XPU deployment benefit from NVLink Fusion’s reduced integration complexity.

Wait before adopting: Workloads primarily consisting of offline training with no real-time inference latency sensitivity retain cost advantage with existing Vera Rubin GPU clusters. Scale-In requires BlueField-4 hardware, necessitating compatibility assessment with existing host platforms.

Final Thoughts

NVIDIA is extending its competitive frontier from individual GPU chips to the complete AI factory architecture. These five technologies are not standalone products but a coordinated system covering computation, networking, storage, and operations—marking AI infrastructure’s transition into a new era of holistic design competition.