The Agentic Post
Breaking
Gemini’s Multimodal Features, Explained  Â·  ChatGPT Custom GPTs, Explained  Â·  What Is Constitutional AI? Explained  Â·  AI Capex Explained for Investors  Â·  AI Startup Valuations: How They Are Set  Â·  How to Reskill for an AI Job Market  ·  
Home/Hardware & Robotics/Chips & GPUs
NVIDIA’s Vera Rubin Chips Ship in Fall

NVIDIA’s Vera Rubin Chips Ship in Fall

Chips & GPUs

NVIDIA's Vera Rubin platform, succeeding Blackwell, entered full production in June 2026 and begins shipping to eight cloud partners this fall, claiming 10x lower inference costs.

NVIDIA’s next-generation chip platform starts shipping to cloud providers this fall. Vera Rubin, the successor to the Blackwell architecture that currently powers most of the industry’s AI training and inference, entered full production earlier this year and is now rolling out to eight named cloud partners, with NVIDIA claiming a 10x reduction in inference token cost over Blackwell.

Quick facts

  • Vera Rubin is a seven-chip, rack-scale AI computing platform succeeding Blackwell, confirmed in full production at NVIDIA’s GTC Taipei keynote on June 1, 2026.
  • Production shipments begin this fall to eight cloud partners: AWS, Azure, Google Cloud, Oracle, CoreWeave, Lambda, Nebius, and Nscale.
  • The flagship Vera Rubin NVL72 rack packs 72 Rubin GPUs and 36 Vera CPUs, delivering a claimed 3.6 exaflops of inference compute in a single liquid-cooled unit.
  • NVIDIA claims 10x lower inference token cost and roughly triple the memory bandwidth per GPU compared to Blackwell, though these figures haven’t been independently verified at production scale.

What’s actually in the box

Vera Rubin isn’t a single chip, it’s seven co-designed components meant to work as one system rather than parts assembled after the fact: the Rubin GPU, the new Vera CPU (replacing Grace), an NVLink 6 switch, a ConnectX-9 SuperNIC, a BlueField-4 DPU, a Spectrum-6 Ethernet switch, and the Groq 3 LPU, added at GTC in March following NVIDIA’s acquisition of Groq. That last addition is specifically aimed at low-latency, deterministic inference for the decode phase of agentic generation, the step where a model actually produces its response token by token, which NVIDIA says pairs with the NVL72 racks to deliver a 35x improvement in inference throughput per megawatt on trillion-parameter models.

Per NVIDIA’s own announcement, each Rubin GPU carries HBM4 memory delivering roughly 22 terabytes per second of bandwidth, close to triple Blackwell’s per-GPU figure, and NVLink 6 doubles rack interconnect speed to 260 terabytes per second, more bandwidth than the entire internet, according to the company. A full Vera Rubin POD scales to 40 racks and 1,152 GPUs for a claimed 60 exaflops of total compute.

The power problem this is actually trying to solve

The more consequential part of the announcement, for anyone building or operating data centers rather than just buying GPUs, is the infrastructure layer NVIDIA is shipping alongside the chips. Working with more than 200 data center infrastructure partners, NVIDIA introduced the DSX platform, including DSX Max-Q for dynamic power provisioning that the company says lets operators deploy 30% more AI infrastructure within a fixed power budget, and DSX Flex, aimed at treating AI factories as grid-flexible assets that could unlock up to 100 gigawatts of power that’s currently effectively stranded on the grid. Power availability, not chip supply, has increasingly become the actual bottleneck on how fast new AI capacity can come online, which is why infrastructure-layer claims like these matter as much as the raw compute specs.

Who’s actually building on it

Beyond the eight cloud partners receiving initial shipments, NVIDIA lists a broad set of AI labs adopting the platform, including Anthropic, Cohere, Meta, Mistral AI, OpenAI, Perplexity, Runway, and xAI, alongside server OEMs Cisco, Dell, HPE, Lenovo, and Supermicro building Vera Rubin systems. Microsoft has specifically committed to deploying Vera Rubin NVL72 racks in its next-generation Fairwater AI superfactory sites. That breadth of adoption across labs that otherwise compete hard against each other says something simple: nearly the entire frontier AI industry is building its next capacity wave on the same underlying hardware platform, whatever their model-level differences.

Common questions

Does this replace Blackwell immediately? No. Blackwell remains the current generation in wide deployment; Vera Rubin is the next generation beginning shipments this fall, and most existing infrastructure will run Blackwell for some time yet.

Can smaller companies access Vera Rubin hardware? Initial shipments go to the eight named cloud partners and their infrastructure; smaller teams will access the platform through those providers’ cloud instances rather than buying hardware directly, similar to how Blackwell access has worked.

What is the Groq 3 LPU doing in an NVIDIA platform? NVIDIA acquired Groq and integrated its low-latency inference chip design directly into Vera Rubin as the seventh co-designed component, specifically to speed up the token-by-token decode phase of agentic AI responses.

Key takeaway

The specific performance multiples NVIDIA is quoting are the company’s own figures, not independently verified benchmarks, so treat 10x and 35x claims as directional until third-party testing catches up. What’s independently checkable is the shipping timeline and partner list, and those point to the same conclusion either way: most of the AI industry’s next round of compute is landing on Vera Rubin hardware starting this fall.

Up Next
Google’s Lyria 3.5 Takes Aim at Suno

Google’s Lyria 3.5 Takes Aim at Suno

Gemini

Google launched Lyria 3.5 inside Flow Music on July 29, 2026, promising more natural vocals, as major labels work toward rules on whether AI-generated music can chart.

Google upgraded its AI music model on July 29, 2026, and the framing from the music industry press was blunt: this is a direct shot at Suno, the AI music startup that’s dominated the category’s attention. Lyria 3.5 promises tracks that sound more natural, with vocals carrying more expression and emotion than its predecessor — and it lands the same week the major record labels are reportedly working out rules for whether AI-generated songs should be allowed on the charts at all.

Quick facts

  • Google rolled out Lyria 3.5 on July 29, 2026, inside Flow Music, its AI music creation platform.
  • The model is pitched specifically on more natural-sounding tracks and more emotionally expressive vocals compared to Lyria 3.
  • Flow Music has a winding history: it began as Riffusion, an open-source project that went viral in 2022, relaunched as ProducerAI in 2025, was acquired by Google in February 2026, and was rebranded Flow Music in April 2026.
  • Believe, a major music distributor, partnered with Google in May 2026 to offer Flow Music to artists across its label and self-release arm, TuneCore.
  • Separately, major record labels are reportedly working toward a framework that would exclude fully AI-generated music from qualifying for official charts.

Where Lyria has already been

Lyria isn’t a new project — Google has been iterating on it in public view for a while. Per Music Business Worldwide’s reporting, an early version of the model powered Dream Track, a YouTube Shorts experiment that let creators generate tracks using AI voice clones of artists including Charlie Puth, T-Pain, and Alec Benjamin. A later version, Lyria 2, powered YouTube’s “Speech to Song” tool, which turns spoken dialogue into music. Lyria 3 followed in February 2026 inside the Gemini app, generating 30-second tracks directly from text prompts or images. Lyria 3.5 is the next step in that same lineage, now living specifically inside the dedicated Flow Music platform rather than as a feature bolted onto Gemini.

Why Google bought a Riffusion-descended startup instead of building from scratch

Flow Music’s origin story matters for understanding what Google actually acquired. Riffusion started as a viral open-source project in 2022, was relaunched commercially as ProducerAI in 2025, and only became a Google product when the company bought the platform in February 2026. That’s a different path than building an in-house consumer music app from zero — Google acquired an existing product, user base, and production tooling, then plugged its own Lyria models in as the underlying generation engine. The Believe and TuneCore distribution partnership, struck in May 2026, extends that further: it gives artists working with a major independent distributor a direct path to release AI-assisted tracks through existing industry channels rather than only through a standalone consumer app.

The bigger fight: does AI-generated music even count?

Underneath the model upgrades, the more consequential story in AI music right now is happening at the industry level. Reporting from AI Music Billboards describes major record labels working toward an agreed framework that would prevent fully AI-generated tracks from qualifying for official music charts, while music made with AI as a production tool under a human songwriter or producer’s direction would continue to qualify. No universal standard has been finalized, but the direction is clear: the industry is moving past debating whether AI belongs in music production at all, and into drawing a specific, practical line between AI-assisted and AI-generated.

That distinction will matter enormously for products like Flow Music. A tool positioned around letting artists direct and refine AI-generated tracks, rather than simply outputting finished songs with no human input, is much better positioned for a chart landscape that draws that exact line — which may be part of why Google is building Lyria into a creative platform with distribution partnerships rather than a pure one-shot generator.

Common questions

Is Lyria 3.5 available to everyone? It’s rolling out inside Flow Music; availability by region and account type wasn’t fully detailed in the reporting reviewed here, so check Flow Music directly for current access in your market.

Can I still use Lyria inside the Gemini app? Lyria 3 launched inside Gemini for quick 30-second generations from prompts or images; Lyria 3.5’s dedicated home is Flow Music, aimed at more complete music creation rather than quick clips.

What counts as “AI-generated” versus “AI-assisted” for chart purposes? No universal standard exists yet. The direction reported so far draws the line at meaningful human creative input: music directed and shaped by a songwriter or producer using AI as a tool is expected to remain chart-eligible, while output with essentially no human creative direction faces growing restriction.

Key takeaway

If you’re evaluating AI music tools for real release plans rather than experimentation, the model quality race (Lyria 3.5 versus Suno versus everyone else) is only half the decision. The other half is whether your workflow keeps enough human creative direction in the loop to stay eligible under whatever chart and licensing rules the industry settles on — a question the model itself can’t answer for you.