
Nvidia reported that its upcoming Vera Rubin NVL72 systems achieve a 30-fold increase in token throughput per megawatt compared to the GB300 NVL72. Analysis using the DeepSeek V4 Pro model indicates these systems reduce the cost per million tokens by up to 45 times through full-stack hardware and software optimization.
Power efficiency and throughput gains
The Vera Rubin NVL72 architecture focuses on maximizing tokens per second within fixed power constraints. Performance data shows these systems produce 30 times more throughput per megawatt than the GB300 NVL72. This efficiency stems from engineering codesign that integrates the model workloads with specific compute, networking, and memory configurations.
By reducing the cost per million tokens by a factor of 45 when running DeepSeek V4 Pro, the system aims to make complex AI workloads more economical. The integration of CUDA-X libraries is intended to maintain hardware productivity across various accelerated workloads after deployment.
Operational constraints and verification
These performance metrics rely on specific configurations and the DeepSeek V4 Pro model. While Nvidia utilizes validated reference designs to standardize deployments, the reported gains assume extreme codesign across the entire stack. Real-world results may vary depending on the specific software optimizations and environmental factors of individual data centers.
The comparison is based on data from SemiAnalysis AgentX. The report does not provide a specific general release date for the Vera Rubin NVL72 systems, focusing instead on the architectural capabilities and thegoverning role of power efficiency in large-scale AI infrastructure.
Original source
This report summarises the source below. Analysis is labelled separately; product and research claims remain attributed to their source.
Read the original at NVIDIA