Skip to content
NextGZ Logo
HomeForum
Content & community
Search everything

One search across products and courses; each area keeps its own experience

Blog

Articles, guides and updates

Live learning
Opening schedule

Live, offline and workshop cohorts

Multi-topic courses
Courses

Chinese, business, marketing, AI, EQ, career skills and Vibe coding

Chinese tools
Learning path
Dictionary
Digital Art
Art courses
My Art courses
AI tools
AI Tutor
Marketplace
Store

Products, memberships and offers

Brushes & Templates
My downloads
About us
Terms & privacy
Login
Log in
Affiliate NextGZ

Biến nội dung học tập của bạn thành thu nhập affiliate

Đăng ký miễn phí, lấy link riêng và giới thiệu NextGZ cho người cần học tiếng Trung hoặc Digital Art.

Hoa hồng 10-25% từ membership, khóa học và resource art
Link riêng, cookie 30 ngày, tracking tự động
Dashboard rõ số liệu, đủ ngưỡng là yêu cầu rút tiền
Đăng ký làm affiliateXem trang cộng tác viên
NextGZ Logo

Học vẽ Digital Art & học tập đa chủ đề với nội dung cập nhật liên tục. Phương pháp học thông minh với AI cho người bận rộn.

Việt Nam 🇻🇳

Học tập

  • Khóa học
  • Luyện phát âm
  • Digital Art
  • Lớp live/offline
  • Tài nguyên Art
  • Cửa hàng
  • AI Tutor

Thông tin

  • Về chúng tôi
  • Liên hệ
  • Dashboard

Đối tác

  • NextGZ
  • GZ Team

Pháp lý

  • Điều khoản
  • Bảo mật
Discord NextGZ Discord avatar

Discord NextGZ

Đang tải
Tham gia
© 2025 NextGZ·Nền tảng học Online: Tiếng Trung, Digital Art, Marketing, AI, Vibe coding, Kỹ Năng EQ cùng nhiều kỹ năng bổ trợ cần thiết khác
Cộng đồng học Digital Art NextGZ
Trang chủBlogForumCá nhân
Đang tải nội dung…
HomeBlogTin Tức AI
Tin Tức AI

NVIDIA announces Rubin GPU architecture: How does HBM4 22 TB/s change AI agents?

NVIDIA announces details of Rubin GPU with 336 billion transistors, 288 GB HBM4, bandwidth 22 TB/s, and Vera Rubin NVL72 rack design for agent AI.

July 22, 2026•6 min read
NVIDIA announces Rubin GPU architecture: How does HBM4 22 TB/s change AI agents?
SCN

Written by

Sen của NextGZ

Admin

Share

Written by

SCN
Sen của NextGZ

Admin

6 min read

Share

In This Article

  • Why did NVIDIA design Rubin specifically for agentic AI?
  • What’s inside the Rubin GPU?
  • HBM4 and 22 TB/s bandwidth: Which bottleneck does it solve?
  • How does Rubin optimize long-context and multi-GPU communication?
  • From GPU to Vera Rubin NVL72: Power becomes a system-level challenge
  • What does Rubin mean for enterprises and AI creators?
  • What remains unconfirmed
Nghe nội dungMiễn phí

1:04 ước tính · Chưa có giọng vi-VN

On July 21, 2026, NVIDIA published a detailed technical analysis of the Rubin GPU architecture, the central component of the Vera Rubin platform for agentic AI systems. The company positions Rubin not just as a step forward in computational power, but as a synchronized design encompassing GPU, HBM4 memory, NVLink, power, cooling, and rack architecture.

A point to read with caution is that most performance specifications currently come from NVIDIA. For example, the claim of “up to 10x agent throughput per unit of power compared to Blackwell” was measured on an internal 2-trillion-parameter MoE workload. This is important primary data for understanding product direction, but it is not yet an independent benchmark in customer deployment environments.

Why did NVIDIA design Rubin specifically for agentic AI?

According to the NVIDIA Technical Blog, agent workloads differ from a single short query–response in that the model must reason through multiple steps, call tools, retrieve data, check intermediate results, and maintain long context. This process simultaneously puts pressure on decode speed, per-step latency, KV cache capacity, memory bandwidth, and inter-GPU communication capabilities.

Therefore, NVIDIA is not only optimizing for peak operations per second. Rubin’s goal is to keep compute units busy throughout the inference chain, minimize data waiting time, reduce gaps between kernels, and maintain throughput when the model is spread across a large GPU domain. This approach reflects a shift from “faster GPU” to “the entire AI factory producing more useful work within the same power envelope.”

What’s inside the Rubin GPU?

NVIDIA Rubin GPU architecture diagram with GPC, HBM4, L2 cache, NVLink, and PCIe Gen 6
NVIDIA Rubin GPU chip architecture diagram. Source: NVIDIA Technical Blog.

NVIDIA announced the Rubin GPU has 336 billion transistors, 224 streaming multiprocessors, and 896 Tensor Cores. Two compute dies are connected within a single package using NVIDIA High-Bandwidth Interface. The third-generation Transformer Engine supports multiple numeric formats and is claimed by the company to achieve up to 50 petaflops NVFP4 for inference.

Compute blocks are organized into Graphics Processor Clusters, coordinating with a centralized L2 cache, GigaThread Engine, MIG Control, and NV-DEC decoder. The goal is to maintain GPU utilization as workloads constantly shift between inference, token generation, retrieval, and tool calling, rather than optimizing for a single fixed kernel type.

Rubin also increases Tensor Core throughput per cycle by processing twice the K size in matrix multiplication. In NVIDIA’s example, a GEMM requiring four K-loop iterations on Blackwell can complete in two iterations on Rubin. The expected benefit is reduced loop overhead and improved performance at large tensor parallel scales; the actual improvement still depends on model configuration and software.

HBM4 and 22 TB/s bandwidth: Which bottleneck does it solve?

Memory bandwidth chart of NVIDIA Hopper, Blackwell, Blackwell Ultra, and Rubin
Memory bandwidth increases to 22 TB/s on Rubin. Source: NVIDIA Technical Blog.

Each Rubin GPU is announced by NVIDIA to support up to 288 GB of HBM4 and a peak bandwidth of 22 TB/s. Compared to the 8 TB/s used in the company’s chart for Blackwell and Blackwell Ultra, this is approximately a 2.8x increase. Larger capacity helps keep models, long contexts, and KV cache closer to the GPU; higher bandwidth helps the token generation phase continuously load weights and states without forcing Tensor Cores to wait too long.

Rubin pairs this memory with an improved Tensor Memory Accelerator. For Mixture of Experts models, a descriptor can be shared among multiple experts and update address fields or strides directly within instructions, instead of maintaining a separate descriptor for each expert. NVIDIA believes this approach reduces metadata management overhead and weight movement overhead as the number of experts scales.

At the interconnect level, NVLink 6 is announced to provide 3,600 GB/s of scale-up bandwidth to the NVLink Switch; NVLink-C2C reaches 1,800 GB/s for synchronous CPU–GPU communication; PCIe Gen 6 x16 achieves up to 256 GB/s on the host side. These figures describe the platform’s technical ceiling, meaning not all applications can utilize the full bandwidth.

How does Rubin optimize long-context and multi-GPU communication?

For long attention, NVIDIA combines 2:4 activation sparsity, adaptive compression, and higher softmax throughput. Intermediate attention data can be converted to a sparse format with metadata, reducing the amount of data to write, store, and process during softmax and subsequent multiplication. The company claims exponential function throughput on Rubin is doubled for FP32 and quadrupled for BF16/FP16 compared to the Blackwell baseline.

Another change is a tile-level dependent kernel scheduling mechanism. A consumer kernel can start earlier as soon as the required data portion is ready, instead of waiting for a large producer group to complete. NVIDIA also adds counted writes for device-initiated NVLink communication, aiming to reduce synchronization steps and keep data moving between GPUs with lower latency.

These improvements are well suited for agent workloads, as each token may pass through multiple sequential dependency stages. However, the impact on per-task cost will still depend on the compiler, libraries, model partitioning, and the ability of software to exploit new features.

From GPU to Vera Rubin NVL72: Power becomes a system-level challenge

NVIDIA DSX MaxLPS chart showing an increase in GPU count within the same power budget
DSX MaxLPS is described by NVIDIA as enabling up to 40% more GPUs within the same power budget. Source: NVIDIA Technical Blog.

At the rack level, the Vera Rubin NVL72 integrates compute, networking, 45°C liquid cooling, power steering, and Intelligent Power Smoothing. NVIDIA states that the new power smoothing mechanism can reduce average power consumption by approximately 10% compared to previous techniques and reduce peak power by about 20% within a 50-millisecond window.

DSX MaxLPS is described by the company as helping to deploy up to 40% more GPUs within the same power budget at an efficient operating point, with minimal impact on workload performance. This figure is a design claim from NVIDIA, not a guarantee that every data center will achieve the same result.

This perspective also helps place Rubin alongside other infrastructure directions. NextGZ has analyzed Microsoft bringing AMD Helios to Azure, where CPU, accelerator, and cloud platform are co-designed for large-scale AI workloads. The competition is therefore no longer limited to individual chips, but extends to networking, racks, power, cooling, and operational software.

What does Rubin mean for enterprises and AI creators?

In the short term, Rubin is most significant for cloud providers, research labs, and enterprises running large models. If NVIDIA’s claims are validated in real-world deployments, the platform could increase the number of concurrently served agents, extend context lengths, and reduce energy per unit of AI work.

For AI creators or small teams, the impact will come indirectly through the pricing, speed, and limits of cloud-based services. There is no basis to conclude that Rubin will immediately make APIs cheaper or models faster for every user. The final cost still depends on infrastructure investment, software, utilization rates, and the commercial policies of each provider.

What remains unconfirmed

The technical article does not provide the selling price of the Rubin GPU, the full cost of the Vera Rubin NVL72, a detailed market-by-market supply schedule, or independent benchmarks on customer workloads. Accessibility through individual clouds, actual configurations, and the level of cost improvement per token also cannot be inferred from architectural specifications alone.

Conclusion: Rubin shows that NVIDIA is optimizing agentic AI at the full-system level: compute, HBM4, NVLink, kernels, power, and rack. The new specifications are very noteworthy, but they require continued verification through third-party benchmarks and deployment data before claims of “up to 10x” or “40% more GPUs” can be translated into business expectations.

Cite This Article

Use the standard citation format when referencing this article.

Worth viewing

NextGZ
Gemini 3.6 Flash Ra Mắt: Giá API, Benchmark, Tính Năng Và So Sánh Toàn Diện
Blog

Gemini 3.6 Flash Ra Mắt: Giá API, Benchmark, Tính Năng Và So Sánh Toàn Diện

10 plugin Codex biến AI viết code thành một nền tảng làm việc thực thụ
Blog

10 plugin Codex biến AI viết code thành một nền tảng làm việc thực thụ

DeepSeek-V4-Flash-0731 chính thức ra mắt: Agent vượt V4-Pro-Preview, hỗ trợ Codex
Blog

DeepSeek-V4-Flash-0731 chính thức ra mắt: Agent vượt V4-Pro-Preview, hỗ trợ Codex

Featured

Gemini 3.6 Flash Ra Mắt: Giá API, Benchmark, Tính Năng Và So Sánh Toàn Diện
Tin Tức AI

Gemini 3.6 Flash Ra Mắt: Giá API, Benchmark, Tính Năng Và So Sánh Toàn Diện

Claude Opus 5 Ra Mắt: Gần Sức Mạnh Fable 5 Với Nửa Giá, Benchmark Và So Sánh Toàn Diện
Tin Tức AI

Claude Opus 5 Ra Mắt: Gần Sức Mạnh Fable 5 Với Nửa Giá, Benchmark Và So Sánh Toàn Diện

DeepSeek-V4-Flash-0731 chính thức ra mắt: Agent vượt V4-Pro-Preview, hỗ trợ Codex
Tin Tức AI

DeepSeek-V4-Flash-0731 chính thức ra mắt: Agent vượt V4-Pro-Preview, hỗ trợ Codex

10 plugin Codex biến AI viết code thành một nền tảng làm việc thực thụ
Công cụ & Phần mềm

10 plugin Codex biến AI viết code thành một nền tảng làm việc thực thụ

View all posts

Article categories

  • Digital Art215
    • Nền Tảng Hội Họa20
    • Phối Cảnh & Background17
    • Màu Sắc & Ánh Sáng16
    • Bắt Đầu Học Digital Art14
    • Bố Cục & Storytelling9
    • Procreate9
    • Anatomy & Figure Drawing6
    • Typography & Phông Chữ6
    • Luyện Tập & Challenge5
    • Adobe Photoshop3
    • Bản Quyền & Sử Dụng Thương Mại3
    • ibisPaint3
    • Portfolio & Nghề Nghiệp3
    • Chất Liệu & Rendering2
  • Công cụ & Phần mềm195
    • 花瓣素材 科技风系列未来感人工智能机器人元素 197657173Tin Tức AI54
    • Vibe Coding24
    • GitHub15
    • Clip Studio Paint11
  • ImagesOpenAI31
    • ChatGPT logo.svgChatGPT11
  • Tutorial vẽ Anime30
  • Digital Painting22
  • 花瓣素材 图标系列3D风蓝色玻璃银行金服务元素 194915847Phân Tích Kinh Doanh22
    • 花瓣素材 装饰系列3D风磨砂玻璃办公文件元素 194990373Tâm Lý Marketing4
    • Huaban 6385841246Tư Duy Kiếm Tiền4
  • Tiếng Trung cho Artist19
  • Tips học tiếng Trung14
Page 1/4

Đánh giá bài viết

More from author

Unity prototype sạch ngay từ ngày đầu: thư mục, assembly boundary và quy tắc đặt tên cho indie team
Unity

Unity prototype sạch ngay từ ngày đầu: thư mục, assembly boundary và quy tắc đặt tên cho indie team

SSen của NextGZ
40
Three.js từ trang trắng đến scene 3D đầu tiên: camera, renderer, ánh sáng và vòng lặp animation

Three.js từ trang trắng đến scene 3D đầu tiên: camera, renderer, ánh sáng và vòng lặp animation

SSen của NextGZ
30
Godot prototype đầu tiên: hiểu Scene, Node và Resource bằng một mini game có thể mở rộng

Godot prototype đầu tiên: hiểu Scene, Node và Resource bằng một mini game có thể mở rộng

SSen của NextGZ
10
Procreate Selection nâng cao: Feather, Save & Load, Color Fill và ColorDrop cho flat màu sạch
Màu Sắc & Ánh Sáng

Procreate Selection nâng cao: Feather, Save & Load, Color Fill và ColorDrop cho flat màu sạch

SSen của NextGZ
200

You might also like

Gemini 3.6 Flash Ra Mắt: Giá API, Benchmark, Tính Năng Và So Sánh Toàn Diện
花瓣素材 科技风系列未来感人工智能机器人元素 197657173Tin Tức AI

Gemini 3.6 Flash Ra Mắt: Giá API, Benchmark, Tính Năng Và So Sánh Toàn Diện

SSen của NextGZ
4940
Claude Opus 5 Ra Mắt: Gần Sức Mạnh Fable 5 Với Nửa Giá, Benchmark Và So Sánh Toàn Diện
花瓣素材 科技风系列未来感人工智能机器人元素 197657173Tin Tức AI

Claude Opus 5 Ra Mắt: Gần Sức Mạnh Fable 5 Với Nửa Giá, Benchmark Và So Sánh Toàn Diện

SSen của NextGZ
3870
DeepSeek-V4-Flash-0731 chính thức ra mắt: Agent vượt V4-Pro-Preview, hỗ trợ Codex
花瓣素材 科技风系列未来感人工智能机器人元素 197657173Tin Tức AI

DeepSeek-V4-Flash-0731 chính thức ra mắt: Agent vượt V4-Pro-Preview, hỗ trợ Codex

SSen của NextGZ
490
10 plugin Codex biến AI viết code thành một nền tảng làm việc thực thụ
Công cụ & Phần mềm

10 plugin Codex biến AI viết code thành một nền tảng làm việc thực thụ

SSen của NextGZ
410

Bình luận

0 bình luận

Đăng nhập để tham gia thảo luận cùng cộng đồng!

Đăng nhập ngay

Đang tải bình luận...

#hạ tầng AI#NVIDIA#Rubin GPU#Vera Rubin#HBM4#NVLink 6#AI tác nhân#Vera Rubin NVL72
Affiliate NextGZ

Biến nội dung học tập của bạn thành thu nhập affiliate

Đăng ký miễn phí, lấy link riêng và giới thiệu NextGZ cho người cần học tiếng Trung hoặc Digital Art.

Hoa hồng 10-25% từ membership, khóa học và resource art
Link riêng, cookie 30 ngày, tracking tự động
Đăng ký làm affiliateXem trang cộng tác viên

Cite This Article

Use the standard citation format when referencing this article.

Worth viewing

NextGZ
Gemini 3.6 Flash Ra Mắt: Giá API, Benchmark, Tính Năng Và So Sánh Toàn Diện
Blog

Gemini 3.6 Flash Ra Mắt: Giá API, Benchmark, Tính Năng Và So Sánh Toàn Diện

10 plugin Codex biến AI viết code thành một nền tảng làm việc thực thụ
Blog

10 plugin Codex biến AI viết code thành một nền tảng làm việc thực thụ

DeepSeek-V4-Flash-0731 chính thức ra mắt: Agent vượt V4-Pro-Preview, hỗ trợ Codex
Blog

DeepSeek-V4-Flash-0731 chính thức ra mắt: Agent vượt V4-Pro-Preview, hỗ trợ Codex

Featured

Gemini 3.6 Flash Ra Mắt: Giá API, Benchmark, Tính Năng Và So Sánh Toàn Diện
Tin Tức AI

Gemini 3.6 Flash Ra Mắt: Giá API, Benchmark, Tính Năng Và So Sánh Toàn Diện

Claude Opus 5 Ra Mắt: Gần Sức Mạnh Fable 5 Với Nửa Giá, Benchmark Và So Sánh Toàn Diện
Tin Tức AI

Claude Opus 5 Ra Mắt: Gần Sức Mạnh Fable 5 Với Nửa Giá, Benchmark Và So Sánh Toàn Diện

DeepSeek-V4-Flash-0731 chính thức ra mắt: Agent vượt V4-Pro-Preview, hỗ trợ Codex
Tin Tức AI

DeepSeek-V4-Flash-0731 chính thức ra mắt: Agent vượt V4-Pro-Preview, hỗ trợ Codex

10 plugin Codex biến AI viết code thành một nền tảng làm việc thực thụ
Công cụ & Phần mềm

10 plugin Codex biến AI viết code thành một nền tảng làm việc thực thụ

View all posts

Article categories

  • Digital Art215
    • Nền Tảng Hội Họa20
    • Phối Cảnh & Background17
    • Màu Sắc & Ánh Sáng16
    • Bắt Đầu Học Digital Art14
    • Bố Cục & Storytelling9
    • Procreate9
    • Anatomy & Figure Drawing6
    • Typography & Phông Chữ6
    • Luyện Tập & Challenge5
    • Adobe Photoshop3
    • Bản Quyền & Sử Dụng Thương Mại3
    • ibisPaint3
    • Portfolio & Nghề Nghiệp3
    • Chất Liệu & Rendering2
  • Công cụ & Phần mềm195
    • 花瓣素材 科技风系列未来感人工智能机器人元素 197657173Tin Tức AI54
    • Vibe Coding24
    • GitHub15
    • Clip Studio Paint11
  • ImagesOpenAI31
    • ChatGPT logo.svgChatGPT11
  • Tutorial vẽ Anime30
  • Digital Painting22
  • 花瓣素材 图标系列3D风蓝色玻璃银行金服务元素 194915847Phân Tích Kinh Doanh22
    • 花瓣素材 装饰系列3D风磨砂玻璃办公文件元素 194990373Tâm Lý Marketing4
    • Huaban 6385841246Tư Duy Kiếm Tiền4
  • Tiếng Trung cho Artist19
  • Tips học tiếng Trung14
Page 1/4

Cộng đồng thực chiến

Zalo - Học AI-Powered Digital Business

Tham gia nhóm để nhận tài nguyên, cập nhật công cụ và trao đổi cách xây dựng digital business cùng AI.

Tham gia nhóm Zalo

Cộng đồng thực chiến

Zalo - Học AI-Powered Digital Business

Tham gia nhóm để nhận tài nguyên, cập nhật công cụ và trao đổi cách xây dựng digital business cùng AI.

Tham gia nhóm Zalo
Dashboard rõ số liệu, đủ ngưỡng là yêu cầu rút tiền

Cộng đồng sáng tạo

Art Bestie Club Discord

Kết nối với cộng đồng Digital Art, chia sẻ tác phẩm và học hỏi quy trình sáng tạo mới.

Tham gia Discord

Cộng đồng sáng tạo

Art Bestie Club Discord

Kết nối với cộng đồng Digital Art, chia sẻ tác phẩm và học hỏi quy trình sáng tạo mới.

Tham gia Discord