Ground Truth.
AI, checked against the source.

Learn · Intermediate

High-bandwidth memory: why AI chips are built around stacks of DRAM

High-bandwidth memory, or HBM, is a kind of computer memory built as a stack of DRAM chips sitting right beside an AI processor, connected by a very wide data path. It exists because modern AI, and text generation in particular, is limited less by how fast a chip can calculate than by how fast it can be fed data. The amount of HBM a chip carries now largely decides which models it can run and how many people it can serve at once.

The memory wall

In 1995, William Wulf and Sally McKee published a short paper called "Hitting the Memory Wall." Their point was simple: processor speed was improving much faster than memory speed, so eventually processors would spend most of their time waiting for data. Three decades later, AI accelerators are the starkest example. A modern GPU can perform an enormous number of multiplications per second, but only if the numbers arrive fast enough.

The roofline model, introduced by Samuel Williams, Andrew Waterman and David Patterson in 2009, makes this precise. For any task, divide the arithmetic it needs by the bytes it must move. If that ratio is low, the task is memory-bound: buying more compute does nothing, and only more bandwidth helps.

Why text generation is memory-bound

When a language model writes text, it produces one token at a time. For each token it must read its active weights from memory, plus the growing KV cache that stores what it has already processed in the conversation. For a single user, that is a lot of reading for relatively little arithmetic. Our lesson on why LLM inference is memory-bound walks through the details, and prefill and decode explains why the token-by-token phase is the bandwidth-hungry one.

A back-of-envelope example: if a model's weights take 70 GB and the memory can deliver about 3.35 terabytes per second, the bandwidth NVIDIA quotes for its H100 SXM, then reading all the weights once takes about a fiftieth of a second. That puts an idealized ceiling of roughly 48 tokens per second on one user's generation, before counting the cache. Serving many users at once helps, because one read of the weights can be shared across a batch, which is why continuous batching matters so much.

How HBM is built

Ordinary graphics memory chips sit around the processor on the circuit board, each with a relatively narrow connection. HBM takes a different approach. Several DRAM dies are stacked vertically and wired together with thousands of tiny vertical connections drilled through the silicon, called through-silicon vias. The stack sits on a silicon base layer next to the processor inside the same package, so the wires between them are extremely short and extremely numerous: through the HBM3E generation, each stack has a 1,024-bit interface.

The analogy is a highway. Graphics memory adds lanes by running faster cars on a few roads; HBM builds a thousand-lane road that is only a few millimeters long. AMD first shipped HBM on a consumer graphics card in 2015, and it has since become standard on data-center AI accelerators. Three companies make nearly all of it: SK hynix, Samsung and Micron.

Capacity matters as much as speed

Bandwidth decides how fast a model runs; capacity decides whether it fits at all. A model's weights, every user's KV cache, and working buffers all have to live in fast memory. That is why accelerators keep adding HBM; NVIDIA's GB300 carries 288 GB. Techniques that shrink memory needs are some of the most important in AI engineering: quantization stores weights in fewer bits, mixture-of-experts models touch only part of their weights per token, and FlashAttention by Tri Dao and colleagues reorganized attention so it moves far less data between HBM and the chip's small on-chip memory.

Why it matters now

HBM has become an economic bottleneck, not just a technical one. This week Epoch AI estimated how many AI agents the memory shipped through 2027 could run, using HBM supply as its yardstick precisely because memory capacity and bandwidth limit how many sessions an accelerator can serve. Memory makers have warned that supply will stay tight into 2027 and 2028, and memory costs show up in the prices of everything from data-center GPUs to desktop AI boxes like the DGX Spark, which uses a different, shared-memory design but faces the same capacity trade-off.

Limits to keep in mind

HBM is expensive and difficult to manufacture, because stacking and connecting dies reduces yields, and its capacity per chip is far below what ordinary server memory offers. That is why systems also use slower, larger memory tiers and offloading for data that does not need to be close at hand. And not every AI workload is memory-bound: processing a long prompt in one go, or training with large batches, can be limited by compute instead.

Key papers
Hitting the Memory Wall: Implications of the Obvious (Wulf and McKee, 1995)
Roofline: An Insightful Visual Performance Model for Multicore Architectures (Williams, Waterman and Patterson, 2009)
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness (Dao et al., 2022)
How many AI agents could run on the AI chips shipped through 2027? (Epoch AI, 2026)

Key questions

What is high-bandwidth memory?

High-bandwidth memory is DRAM built as a vertical stack of memory dies connected by tiny vertical wires and placed right next to the processor, which gives it a very wide connection and far more bandwidth than ordinary graphics memory.

Why does AI inference need so much memory bandwidth?

Generating each new token requires reading the model's active weights and the conversation's cached state from memory, so the speed at which memory can deliver data, not raw arithmetic, usually limits how fast text comes out.

How is HBM different from the RAM in a laptop?

Laptop RAM sits on separate modules connected by a relatively narrow bus, while HBM is packaged beside the processor with an interface over a thousand bits wide per stack, trading cost and capacity for much higher bandwidth.
Cite this

APA

Ground Truth. (2026, October 3). High-bandwidth memory: why AI chips are built around stacks of DRAM. Ground Truth. https://groundtruth.day/learn/high-bandwidth-memory.html

BibTeX

@misc{groundtruth:high-bandwidth-memory,
  title  = {High-bandwidth memory: why AI chips are built around stacks of DRAM},
  author = {{Ground Truth}},
  year   = {2026},
  month  = {oct},
  url    = {https://groundtruth.day/learn/high-bandwidth-memory.html}
}

Topics: hardware · memory · inference · infrastructure · fundamentals