All stories
AI

SanDisk and SK Hynix Unveil HBF Spec for Terabyte-Scale GPU Memory

SanDisk and SK Hynix have formally introduced the Hybrid Bonding Fabric (HBF) specification, a groundbreaking initiative set to provide GPUs with terabytes of supplementary memory and an eventual bandwidth reaching 3 TB/s.

By TECH NEWS Editorial·Source:Tom's Hardware·4 min read·2h ago

This content was summarized and interpreted by AI; it may contain errors — please verify accuracy with the original sources. Learn more

Share

Listen to this story

0:00 / 0:00
SanDisk and SK Hynix Unveil HBF Spec for Terabyte-Scale GPU Memory

SanDisk and SK Hynix have formally introduced the Hybrid Bonding Fabric (HBF) specification, a groundbreaking initiative poised to equip GPUs with terabytes of supplementary memory and an eventual bandwidth reaching 3 TB/s. This collaborative effort leverages advanced 16-Hi NAND stacks and the open Universal Chiplet Interconnect Express (UCIe) standard, fundamentally altering the landscape of GPU memory architecture. While the immediate impact remains to be seen, with only four companies reportedly expressing initial interest, HBF represents a significant leap towards addressing the escalating memory demands of artificial intelligence (AI) and high-performance computing (HPC) workloads.

The significance of HBF lies in its audacious promise of vastly expanded memory capacity, a critical bottleneck for contemporary GPUs, particularly in the era of colossal AI models. Unlike High Bandwidth Memory (HBM), which prioritizes extreme speed through vertically stacked DRAM dies, HBF aims to provide massive, cost-effective capacity by integrating high-density NAND flash. Current top-tier GPUs, such as NVIDIA's H200, feature up to 141GB of HBM3e, delivering over 4.8 TB/s of bandwidth. While HBM excels at feeding the GPU's processing units with data at breakneck speeds, its capacity remains limited and expensive due to DRAM's inherent cost and density constraints. HBF seeks to complement, rather than replace, HBM by offering a secondary tier of high-capacity memory, allowing GPUs to handle datasets that currently necessitate offloading to slower, CPU-attached system memory or even network-attached storage. This approach could drastically reduce the time and energy spent on data transfers, accelerating training for large language models (LLMs) and complex simulations that regularly exceed the gigabyte capacities of even the most advanced HBM configurations.

The technical underpinnings of HBF are crucial to its potential impact. The specification's reliance on 16-Hi NAND stacks signifies a dramatic increase in vertical integration, enabling unprecedented memory density within a compact footprint. This stacking is facilitated by hybrid bonding technology, a direct wafer-to-wafer or die-to-wafer bonding technique that creates extremely fine-pitch electrical connections, vastly improving data transfer efficiency and reducing latency compared to traditional wire bonding. The integration of UCIe is equally pivotal. UCIe provides an open industry standard for chiplet interconnects, allowing different types of dies—such as a GPU die, HBM stacks, and now HBF NAND stacks—to communicate efficiently within a single package. This modularity offers unprecedented flexibility for GPU designers, enabling them to customize memory configurations based on specific application needs. For instance, an AI training accelerator could feature both high-speed HBM for immediate data access and multi-terabyte HBF for storing vast model parameters or training datasets, all interconnected seamlessly via UCIe. This capability could democratize access to larger models by lowering the hardware entry barrier, as NAND is significantly cheaper per bit than DRAM.

Compared to existing solutions, HBF offers a distinct value proposition. Traditional GPU memory, primarily GDDR, provides high bandwidth but is physically limited by the PCB traces and controller complexity when scaling capacity. HBM solves the bandwidth problem through 3D stacking but remains capacity-limited and expensive. While some solutions like CXL (Compute Express Link) allow CPUs and GPUs to share memory pools, they typically involve traversing PCIe-based interconnects, introducing latency. HBF, by integrating NAND directly into the GPU package via hybrid bonding and UCIe, aims for much lower latency and higher bandwidth than external memory solutions, while offering capacities far beyond what HBM can economically provide. This creates a new memory tier, bridging the gap between ultra-fast, limited-capacity HBM and slower, high-capacity system memory.

Looking ahead, HBF's success hinges on broader industry adoption and the development of robust software ecosystems. The current interest from only four companies suggests that while the technology is promising, it faces the typical chicken-and-egg challenge of new standards: hardware needs software support, and software developers need compelling hardware. If HBF gains traction, it could lead to the emergence of "memory-rich" GPUs, fundamentally altering how AI models are designed and deployed. Data centers could see significant power savings by reducing the need for extensive data shuffling between different memory tiers, and researchers could explore even larger, more complex models without hitting memory capacity walls. The eventual 3 TB/s bandwidth, combined with terabytes of capacity, could enable a new class of AI accelerators that are both powerful and cost-efficient. However, challenges remain in optimizing NAND for frequent random access patterns typical of GPU workloads, and ensuring long-term reliability and endurance. The next few years will be critical in determining whether HBF evolves from a niche solution to a foundational technology shaping the future of high-performance computing and AI.