NVIDIA Unveils Five Core Technologies for Vera Rubin AI Factory Ecosystem

On August 24, 2024 (local time), at the Hot Chips conference, NVIDIA announced major advancements in the Vera Rubin AI factory ecosystem:
- NVIDIA Groq 3 LPX: Entered full-volume production as a low-latency inference accelerator for token generation
- Spectrum-X Multi-Plane: Enables scaling to 512,000 GPUs without adding Layer-3 networking
- Scale-In: Based on BlueField-4 and DOCA, handles security, storage access, and operations offload
- NVLink Fusion: Enables third-party XPU and CPU integration into NVIDIA’s rack architecture
Initial Deployment: Nebius becomes the first adopter, integrating Groq 3 LPX into its production inference platform Nebius Token Factory, allowing developers to retain their existing API stack without migration.
Groq 3 LPX: Low-Latency Token Generation Accelerator
Groq 3 LPX is explicitly defined by NVIDIA as an “interactive AI inference accelerator”—not a replacement for Vera Rubin GPU but as an extension that specifically accelerates token generation phase.
Applications like code agents involve hundreds to thousands of execution steps, where per-generation latency accumulates, directly impacting end-to-end task completion time. Test results show:
- Output rate of 3,400 tokens/second on Gemma 4 31B with 100K token context in Artificial Analysis benchmark, the fastest recorded for this model
- Response speed up to 4x faster than closest alternative platforms on agent and latency-sensitive workloads
Deployment will prioritize AI cloud providers, with Nebius first and Groq also planned as an early adopter.
Spectrum-X Multi-Plane: Scaling to 512K GPUs Without Layer-3 Overhead

Legacy two-layer Ethernet clusters often require adding Layer-3 networking whenscaling beyond certain sizes, incurring additional hops, latency, jitter, and hardware costs.
The surprising achievement of Spectrum-X multi-plane: In an eight-plane topology, one plane failure still retains ~90% total bandwidth, with hardware-based recovery 11x faster than software alternatives, resulting in 1.6x AI factory output improvement.
Key technical specifications:
- Spectrum-X SN6000 switch uses 102.4Tb/s Spectrum-6 ASIC
- ConnectX-9 SuperNIC delivers up to 1,600Gb/s bandwidth per GPU
- Hardware engine in SuperNIC handles traffic allocation and failover
- Spectrum-XGS connects multiple data centers, improving multi-site NCCL collective performance by 1.9x
Scale-In and NVLink Fusion: Security Offload and Open Integration
Scale-In is dubbed the “fifth pillar” of NVIDIA’s AI networking architecture. Powered by BlueField-4 processor and DOCA software stack, it isolates multi-tenant networking, storage access, security, resource provisioning, and real-time observability into a dedicated hardware domain, preventing continuous CPU/GPU resource consumption.
Complementing this is NVLink Fusion, which extends NVLink beyond NVIDIA GPU interconnect to third-party XPU and CPU via sixth-generation NVLink, NVLink Switch, and NVLink-C2C:
- Sixth-generation NVLink supports 72-XPU interconnect domains with endpoint-to-endpoint latency reduced to one-third and packet rate increased 10x versus general Ethernet solutions
- NVLink-C2C connects XPU to NVIDIA Vera CPU or ecosystem CPUs with up to 6x power efficiency over PCIe
This enables hyperscalers to deploy custom XPU alongside NVIDIA GPUs in existing MGX rack infrastructure (power, cooling, networking, management software, supply chain), allowing dynamic chip mix adjustments based on supply and workload.
Key Technical Comparison

| Technology | Core Function | Key Metric | Baseline | Improvement |
|---|---|---|---|---|
| Groq 3 LPX | Token generation acceleration | 3,400 token/s (Gemma 4 31B) | Closest alternative | 4x faster response |
| Spectrum-X multi-plane | Large-scale cluster scaling | 512K GPUs without Layer-3 | Traditional 3-layer | 11x faster recovery, 1.6x output |
| NVLink Fusion | XPU integration | 72-XPU interconnect domain | Ethernet-based | 1/3 latency, 10x packet rate |
| Scale-In | Infrastructure service offload | BlueField-4 + DOCA | Host CPU/GPU | Isolates security/storage/Ops |
##落地建议 (Application Guidance)
Suitable for immediate adoption: Cloud providers or enterprise AI platforms with existing large-scale GPU clusters facing token generation latency or network scaling bottlenecks should evaluate the Groq 3 LPX plus Spectrum-X multi-plane combination. Hyperscalers already using MGX racks and planning custom XPU deployment benefit from NVLink Fusion’s reduced integration complexity.
Wait before adopting: Workloads primarily consisting of offline training with no real-time inference latency sensitivity retain cost advantage with existing Vera Rubin GPU clusters. Scale-In requires BlueField-4 hardware, necessitating compatibility assessment with existing host platforms.
Final Thoughts
NVIDIA is extending its competitive frontier from individual GPU chips to the complete AI factory architecture. These five technologies are not standalone products but a coordinated system covering computation, networking, storage, and operations—marking AI infrastructure’s transition into a new era of holistic design competition.
