Image: vllm.ai · rights & removal
vLLM Support for NVIDIA Vera Rubin NVL72: 7.8x Throughput over GB200 NVL72
Reporting by vLLM BlogRead the original at vllm.ai
Executive Summary
Facts Only
* vLLM supports Vera Rubin NVL72 hardware.
* Vera Rubin NVL72 has 5x the NVFP4 FLOPS of GB200 NVL72 and 2.4x higher HBM bandwidth.
* The platform features sixth-generation NVLink delivering up to 1.7x more bandwidth than Blackwell.
* vLLM incorporates Rubin-tuned attention, GEMM, and MoE kernels via FlashInfer 0.7.0.
* Locality domains in CUDA 13.4 are leveraged to split MoE weights, optimizing memory access for MoE decode.
* Early results show 7.8x throughput per GPU compared to GB200 NVL72 on AgentX at matched interactivity.
* MLPerf Inference v6.1 testing showed up to 3.7x higher VLM throughput than GB300 NVL72.
* vLLM supports models including DeepSeek, Kimi, GLM, and MiniMax on Rubin.
* Users can deploy nightly images with CUDA 13.4 and PyTorch 2.15 for Rubin hardware.
Full Take
From the original · vLLM Blog
vLLM now supports Vera Rubin NVL72! NVIDIA Vera Rubin is the next-generation platform built for agentic inference.Read the full story at vllm.ai
Sentinel — provisional
No strong signs of machine writing were found in the source article. Provisional estimate, not a finding that a person wrote it.
This text appears to be a human-authored summary or technical post detailing collaborative advancements in accelerating LLM inference on new NVIDIA hardware, supported by community development efforts.
This looks only at the wording of the original source article, not at this page's AI-written sections. A small local AI model made this estimate. It has not been checked against known human and machine texts, so treat it as provisional. It cannot show who wrote an article.
