NVIDIA’s next-generation chip platform starts shipping to cloud providers this fall. Vera Rubin, the successor to the Blackwell architecture that currently powers most of the industry’s AI training and inference, entered full production earlier this year and is now rolling out to eight named cloud partners, with NVIDIA claiming a 10x reduction in inference token cost over Blackwell.
Quick facts
- Vera Rubin is a seven-chip, rack-scale AI computing platform succeeding Blackwell, confirmed in full production at NVIDIA’s GTC Taipei keynote on June 1, 2026.
- Production shipments begin this fall to eight cloud partners: AWS, Azure, Google Cloud, Oracle, CoreWeave, Lambda, Nebius, and Nscale.
- The flagship Vera Rubin NVL72 rack packs 72 Rubin GPUs and 36 Vera CPUs, delivering a claimed 3.6 exaflops of inference compute in a single liquid-cooled unit.
- NVIDIA claims 10x lower inference token cost and roughly triple the memory bandwidth per GPU compared to Blackwell, though these figures haven’t been independently verified at production scale.
What’s actually in the box
Vera Rubin isn’t a single chip, it’s seven co-designed components meant to work as one system rather than parts assembled after the fact: the Rubin GPU, the new Vera CPU (replacing Grace), an NVLink 6 switch, a ConnectX-9 SuperNIC, a BlueField-4 DPU, a Spectrum-6 Ethernet switch, and the Groq 3 LPU, added at GTC in March following NVIDIA’s acquisition of Groq. That last addition is specifically aimed at low-latency, deterministic inference for the decode phase of agentic generation, the step where a model actually produces its response token by token, which NVIDIA says pairs with the NVL72 racks to deliver a 35x improvement in inference throughput per megawatt on trillion-parameter models.
Per NVIDIA’s own announcement, each Rubin GPU carries HBM4 memory delivering roughly 22 terabytes per second of bandwidth, close to triple Blackwell’s per-GPU figure, and NVLink 6 doubles rack interconnect speed to 260 terabytes per second, more bandwidth than the entire internet, according to the company. A full Vera Rubin POD scales to 40 racks and 1,152 GPUs for a claimed 60 exaflops of total compute.
The power problem this is actually trying to solve
The more consequential part of the announcement, for anyone building or operating data centers rather than just buying GPUs, is the infrastructure layer NVIDIA is shipping alongside the chips. Working with more than 200 data center infrastructure partners, NVIDIA introduced the DSX platform, including DSX Max-Q for dynamic power provisioning that the company says lets operators deploy 30% more AI infrastructure within a fixed power budget, and DSX Flex, aimed at treating AI factories as grid-flexible assets that could unlock up to 100 gigawatts of power that’s currently effectively stranded on the grid. Power availability, not chip supply, has increasingly become the actual bottleneck on how fast new AI capacity can come online, which is why infrastructure-layer claims like these matter as much as the raw compute specs.
Who’s actually building on it
Beyond the eight cloud partners receiving initial shipments, NVIDIA lists a broad set of AI labs adopting the platform, including Anthropic, Cohere, Meta, Mistral AI, OpenAI, Perplexity, Runway, and xAI, alongside server OEMs Cisco, Dell, HPE, Lenovo, and Supermicro building Vera Rubin systems. Microsoft has specifically committed to deploying Vera Rubin NVL72 racks in its next-generation Fairwater AI superfactory sites. That breadth of adoption across labs that otherwise compete hard against each other says something simple: nearly the entire frontier AI industry is building its next capacity wave on the same underlying hardware platform, whatever their model-level differences.
Common questions
Does this replace Blackwell immediately? No. Blackwell remains the current generation in wide deployment; Vera Rubin is the next generation beginning shipments this fall, and most existing infrastructure will run Blackwell for some time yet.
Can smaller companies access Vera Rubin hardware? Initial shipments go to the eight named cloud partners and their infrastructure; smaller teams will access the platform through those providers’ cloud instances rather than buying hardware directly, similar to how Blackwell access has worked.
What is the Groq 3 LPU doing in an NVIDIA platform? NVIDIA acquired Groq and integrated its low-latency inference chip design directly into Vera Rubin as the seventh co-designed component, specifically to speed up the token-by-token decode phase of agentic AI responses.
Key takeaway
The specific performance multiples NVIDIA is quoting are the company’s own figures, not independently verified benchmarks, so treat 10x and 35x claims as directional until third-party testing catches up. What’s independently checkable is the shipping timeline and partner list, and those point to the same conclusion either way: most of the AI industry’s next round of compute is landing on Vera Rubin hardware starting this fall.

Leave a Reply