NVIDIA NVHBM Targets Faster, Denser Custom AI Accelerators with NVLink Fusion
NVIDIA says its NVHBM base-die technology can give custom AI accelerators more memory bandwidth, more usable compute area and lower HBM power inside NVLink Fusion systems.
NVIDIA extends NVLink Fusion down to the memory stack
NVIDIA introduced new details for NVHBM on August 26, 2026, positioning the custom HBM base-die technology as a companion to NVLink Fusion for hyperscalers and AI-native companies building their own AI accelerators. The goal is to make custom XPUs easier to integrate into NVIDIA's rack-scale infrastructure while improving the memory bandwidth, silicon-area efficiency and power profile of each accelerator package.
NVLink Fusion is NVIDIA's IP and interconnect framework for connecting custom XPUs and CPUs into the broader NVIDIA AI infrastructure stack. NVHBM addresses a different bottleneck: the local memory interface between a custom accelerator and its HBM stacks.
NVIDIA's headline NVHBM claims
Compared with standard HBM4e, NVIDIA says NVHBM can provide up to 30% more memory bandwidth per stack, up to 25% more compute-die area for accelerator features, and up to 15% lower HBM power usage. NVIDIA further says that combining these gains can translate to roughly 30% higher end-to-end XPU performance in its modeled platform scenarios. These are vendor-reported figures rather than independent benchmark results and will depend on the final XPU design and workload.
NVIDIA says the area savings come from a custom base die and physical interface that move the memory controller into the 3D HBM stack and reduce I/O overhead. Its technical post reports up to a 67% reduction in PHY and support area and up to 80% more usable silicon across the package layout, giving XPU designers more room for matrix engines, SRAM, cache or workload-specific logic.
Why memory matters more for large-model inference
Modern AI accelerators frequently spend substantial time moving model weights, KV-cache data and activations rather than performing arithmetic. That makes memory bandwidth and memory power direct constraints on token throughput and serving density, especially for large language models and long-context workloads.
NVIDIA argues that higher per-stack bandwidth helps keep compute engines fed during memory-bound inference, while lower HBM power creates thermal and electrical headroom that can be redirected to additional compute. At data-center scale, the company says those savings can compound across thousands of accelerators.
NVHBM plus NVLink Fusion at rack scale
Local HBM improvements do not solve distributed model execution by themselves. NVLink Fusion is intended to connect custom XPUs through NVIDIA's sixth-generation NVLink fabric so many accelerators can operate as one scale-up domain. This is particularly relevant for mixture-of-experts serving, where expert parallelism can require frequent movement of activations and hidden states across devices.
The NVLink Fusion chiplet bridges a custom XPU into the scale-up fabric, while NVLink-C2C can connect accelerator and CPU components upstream. NVIDIA's pitch is that a hyperscaler can keep its own custom compute architecture while using NVIDIA networking, rack design and software around it.
A strategic move toward semi-custom AI factories
The announcement matters because custom AI silicon is becoming a larger part of hyperscaler strategy. Instead of choosing between a fully proprietary stack and a standard GPU platform, NVLink Fusion and NVHBM create a semi-custom path: companies can design their own XPU while adopting NVIDIA's memory base-die technology and scale-up infrastructure.
That could shorten integration and qualification work, but it also deepens dependency on NVIDIA's interconnect ecosystem. For infrastructure teams, the real questions will be workload-level performance, memory-vendor support, package economics, power density, software compatibility and how easily capacity can be reprovisioned between custom XPUs and GPUs.
NVHBM is an infrastructure technology announcement, not a new standalone GPU or general-purpose model. The reported performance and power figures should be treated as NVIDIA engineering claims until validated in shipping partner systems.
This article is built from the source material below. Open the originals for full context and the latest updates.