NVIDIA announces Rubin GPU architecture: How does HBM4 22 TB/s change AI agents?
NVIDIA announces details of Rubin GPU with 336 billion transistors, 288 GB HBM4, bandwidth 22 TB/s, and Vera Rubin NVL72 rack design for agent AI.

1:04 ước tính · Chưa có giọng vi-VN
On July 21, 2026, NVIDIA published a detailed technical analysis of the Rubin GPU architecture, the central component of the Vera Rubin platform for agentic AI systems. The company positions Rubin not just as a step forward in computational power, but as a synchronized design encompassing GPU, HBM4 memory, NVLink, power, cooling, and rack architecture.
A point to read with caution is that most performance specifications currently come from NVIDIA. For example, the claim of “up to 10x agent throughput per unit of power compared to Blackwell” was measured on an internal 2-trillion-parameter MoE workload. This is important primary data for understanding product direction, but it is not yet an independent benchmark in customer deployment environments.
Why did NVIDIA design Rubin specifically for agentic AI?
According to the NVIDIA Technical Blog, agent workloads differ from a single short query–response in that the model must reason through multiple steps, call tools, retrieve data, check intermediate results, and maintain long context. This process simultaneously puts pressure on decode speed, per-step latency, KV cache capacity, memory bandwidth, and inter-GPU communication capabilities.
Therefore, NVIDIA is not only optimizing for peak operations per second. Rubin’s goal is to keep compute units busy throughout the inference chain, minimize data waiting time, reduce gaps between kernels, and maintain throughput when the model is spread across a large GPU domain. This approach reflects a shift from “faster GPU” to “the entire AI factory producing more useful work within the same power envelope.”
What’s inside the Rubin GPU?

NVIDIA announced the Rubin GPU has 336 billion transistors, 224 streaming multiprocessors, and 896 Tensor Cores. Two compute dies are connected within a single package using NVIDIA High-Bandwidth Interface. The third-generation Transformer Engine supports multiple numeric formats and is claimed by the company to achieve up to 50 petaflops NVFP4 for inference.
Compute blocks are organized into Graphics Processor Clusters, coordinating with a centralized L2 cache, GigaThread Engine, MIG Control, and NV-DEC decoder. The goal is to maintain GPU utilization as workloads constantly shift between inference, token generation, retrieval, and tool calling, rather than optimizing for a single fixed kernel type.
Rubin also increases Tensor Core throughput per cycle by processing twice the K size in matrix multiplication. In NVIDIA’s example, a GEMM requiring four K-loop iterations on Blackwell can complete in two iterations on Rubin. The expected benefit is reduced loop overhead and improved performance at large tensor parallel scales; the actual improvement still depends on model configuration and software.
HBM4 and 22 TB/s bandwidth: Which bottleneck does it solve?

Each Rubin GPU is announced by NVIDIA to support up to 288 GB of HBM4 and a peak bandwidth of 22 TB/s. Compared to the 8 TB/s used in the company’s chart for Blackwell and Blackwell Ultra, this is approximately a 2.8x increase. Larger capacity helps keep models, long contexts, and KV cache closer to the GPU; higher bandwidth helps the token generation phase continuously load weights and states without forcing Tensor Cores to wait too long.
Rubin pairs this memory with an improved Tensor Memory Accelerator. For Mixture of Experts models, a descriptor can be shared among multiple experts and update address fields or strides directly within instructions, instead of maintaining a separate descriptor for each expert. NVIDIA believes this approach reduces metadata management overhead and weight movement overhead as the number of experts scales.
At the interconnect level, NVLink 6 is announced to provide 3,600 GB/s of scale-up bandwidth to the NVLink Switch; NVLink-C2C reaches 1,800 GB/s for synchronous CPU–GPU communication; PCIe Gen 6 x16 achieves up to 256 GB/s on the host side. These figures describe the platform’s technical ceiling, meaning not all applications can utilize the full bandwidth.
How does Rubin optimize long-context and multi-GPU communication?
For long attention, NVIDIA combines 2:4 activation sparsity, adaptive compression, and higher softmax throughput. Intermediate attention data can be converted to a sparse format with metadata, reducing the amount of data to write, store, and process during softmax and subsequent multiplication. The company claims exponential function throughput on Rubin is doubled for FP32 and quadrupled for BF16/FP16 compared to the Blackwell baseline.
Another change is a tile-level dependent kernel scheduling mechanism. A consumer kernel can start earlier as soon as the required data portion is ready, instead of waiting for a large producer group to complete. NVIDIA also adds counted writes for device-initiated NVLink communication, aiming to reduce synchronization steps and keep data moving between GPUs with lower latency.
These improvements are well suited for agent workloads, as each token may pass through multiple sequential dependency stages. However, the impact on per-task cost will still depend on the compiler, libraries, model partitioning, and the ability of software to exploit new features.
From GPU to Vera Rubin NVL72: Power becomes a system-level challenge

At the rack level, the Vera Rubin NVL72 integrates compute, networking, 45°C liquid cooling, power steering, and Intelligent Power Smoothing. NVIDIA states that the new power smoothing mechanism can reduce average power consumption by approximately 10% compared to previous techniques and reduce peak power by about 20% within a 50-millisecond window.
DSX MaxLPS is described by the company as helping to deploy up to 40% more GPUs within the same power budget at an efficient operating point, with minimal impact on workload performance. This figure is a design claim from NVIDIA, not a guarantee that every data center will achieve the same result.
This perspective also helps place Rubin alongside other infrastructure directions. NextGZ has analyzed Microsoft bringing AMD Helios to Azure, where CPU, accelerator, and cloud platform are co-designed for large-scale AI workloads. The competition is therefore no longer limited to individual chips, but extends to networking, racks, power, cooling, and operational software.
What does Rubin mean for enterprises and AI creators?
In the short term, Rubin is most significant for cloud providers, research labs, and enterprises running large models. If NVIDIA’s claims are validated in real-world deployments, the platform could increase the number of concurrently served agents, extend context lengths, and reduce energy per unit of AI work.
For AI creators or small teams, the impact will come indirectly through the pricing, speed, and limits of cloud-based services. There is no basis to conclude that Rubin will immediately make APIs cheaper or models faster for every user. The final cost still depends on infrastructure investment, software, utilization rates, and the commercial policies of each provider.
What remains unconfirmed
The technical article does not provide the selling price of the Rubin GPU, the full cost of the Vera Rubin NVL72, a detailed market-by-market supply schedule, or independent benchmarks on customer workloads. Accessibility through individual clouds, actual configurations, and the level of cost improvement per token also cannot be inferred from architectural specifications alone.
Conclusion: Rubin shows that NVIDIA is optimizing agentic AI at the full-system level: compute, HBM4, NVLink, kernels, power, and rack. The new specifications are very noteworthy, but they require continued verification through third-party benchmarks and deployment data before claims of “up to 10x” or “40% more GPUs” can be translated into business expectations.
Đánh giá bài viết
More from author

Hướng dẫn tạo và quản lý Ruler trong Clip Studio Paint: Guide Line, Symmetry và dựng hình Manga

Krita dựng phối cảnh thực dụng cho background anime: đường chân trời, điểm tụ và kiểm tra sai lệch

OpenAI ra mắt chương trình ChatGPT cho doanh nghiệp nhỏ: Có gì đáng chú ý?

Gemini 3.6 Flash Ra Mắt: Giá API, Benchmark, Tính Năng Và So Sánh Toàn Diện
You might also like

Gemini 3.6 Flash Ra Mắt: Giá API, Benchmark, Tính Năng Và So Sánh Toàn Diện

10 plugin Codex biến AI viết code thành một nền tảng làm việc thực thụ

Grok 4.5 và GPT-5.6 thực sự đứng đâu khi so với Claude Opus?

Bình luận
0 bình luận
Đăng nhập để tham gia thảo luận cùng cộng đồng!
Đăng nhập ngayĐang tải bình luận...