ИИ-инфраструктура / ИИ-брифинг

Nvidia Vera Rubin architecture increases throughput for DeepSeek V4 Pro workloads

Nvidia Vera Rubin systems deliver 30x higher power efficiency and 45x lower token costs for DeepSeek V4 Pro models compared to previous Blackwell generations.

Image accompanying the original report at NVIDIA
Из NVIDIA. Оригинальное изображение из источника.

Nvidia reported that its upcoming Vera Rubin NVL72 systems achieve a 30-fold increase in token throughput per megawatt compared to the GB300 NVL72. Analysis using the DeepSeek V4 Pro model indicates these systems reduce the cost per million tokens by up to 45 times through full-stack hardware and software optimization.

Power efficiency and throughput gains

The Vera Rubin NVL72 architecture focuses on maximizing tokens per second within fixed power constraints. Performance data shows these systems produce 30 times more throughput per megawatt than the GB300 NVL72. This efficiency stems from engineering codesign that integrates the model workloads with specific compute, networking, and memory configurations.

By reducing the cost per million tokens by a factor of 45 when running DeepSeek V4 Pro, the system aims to make complex AI workloads more economical. The integration of CUDA-X libraries is intended to maintain hardware productivity across various accelerated workloads after deployment.

Operational constraints and verification

These performance metrics rely on specific configurations and the DeepSeek V4 Pro model. While Nvidia utilizes validated reference designs to standardize deployments, the reported gains assume extreme codesign across the entire stack. Real-world results may vary depending on the specific software optimizations and environmental factors of individual data centers.

The comparison is based on data from SemiAnalysis AgentX. The report does not provide a specific general release date for the Vera Rubin NVL72 systems, focusing instead on the architectural capabilities and thegoverning role of power efficiency in large-scale AI infrastructure.

Первоисточник

В этом отчете кратко изложен материал указанного ниже источника. Аналитика приводится отдельно; заявления о продуктах и исследованиях атрибутируются их авторам.

Читать оригинал на NVIDIA

← Назад ко всем новостям ИИ