Silicon Almanac.All explainers
Drag the scene to turn it · scroll to continue

Advanced packaging and HBM

Chips on chips.

An AI accelerator is not one chip. It is a compute die surrounded by towers of stacked memory, all standing on a slab of silicon wiring, all on one package.

Shrinking transistors made chips faster. Packaging them closer together is what keeps them fed. This page takes one of these packages apart.

2 TB/sfrom one HBM4 stack
16DRAM dies in one stack
775 µmheight limit, all of it

HBM4 figures are from the JEDEC standard, JESD270-4.

01 The memory wall

Compute is starving.

A modern AI chip can calculate far faster than ordinary memory can feed it. The fix is not a faster road but a wider one, placed right next door.

A DDR5 memory channel is 64 bits wide and runs centimetres across a circuit board. One HBM4 stack talks over 2,048 bits, a few millimetres from the chip it feeds.

DDR5-6400, one channel51 GB/s
HBM4, one stack2,000 GB/s

DDR5 figure is arithmetic: 6,400 million transfers a second on 8 bytes. Lines are drawn in groups: the amber bus shows 8 of its 64 bits, the cyan one 256 of its 2,048.

02 High Bandwidth Memory

A skyscraper of DRAM.

HBM stacks memory dies on top of a logic die, and wires them together vertically, straight through the silicon. JEDEC's HBM4 standard allows 4, 8, 12 or 16 dies, up to 64 GB per stack.

The catch is height. The whole tower has to fit in 775 µm. That is the thickness of one ordinary, unthinned 300 mm wafer, for up to seventeen dies.

12-high
Capacity, 32 Gb dies48 GB
Height budget per die≈ 59 µm

Semiconductor Engineering reports DRAM dies at 30 to 50 µm today. JEDEC raised the limit from 720 µm to 775 µm for HBM4. The per-die budget is arithmetic and includes the bonds between dies.

03 Through-silicon vias

Copper roads through the chip.

Normal chips only have wires on top. To stack them, every die needs vertical connections that go straight through its silicon: through-silicon vias. They are drilled with the same deep etch as a memory hole, filled with copper, and exposed by grinding the die thin from the back.

04 Joining dies

Solder, or no solder at all.

Stacked dies are joined through tiny solder bumps. Solder needs room, so bumps cannot sit much closer than tens of micrometres. Hybrid bonding drops the solder: two polished dies are pressed together and their copper pads fuse directly.

40 µm
Pads per mm²625
JointSolder microbump

Microbumps have been around 40 µm, moving towards 10 µm for HBM4, per Semiconductor Engineering. TSMC's hybrid bonding went from 9 µm in 2023 to 6 µm in 2025, with 4.5 µm planned for 2029, per Tom's Hardware. Because JEDEC raised the height limit, HBM4 kept microbumps; hybrid-bonded HBM is expected with HBM4E or HBM5. Pads per mm² is arithmetic on the pitch.

05 The interposer

Bigger than a chip can be.

A scanner can print at most one reticle field, about 858 mm². The dies here are wired together through a silicon interposer, which is made by stitching several fields together. TSMC calls its version CoWoS, and it keeps growing.

Interposer area4,719 mm²
HBM stacks12
When2025–2026
Package substrateover 100 × 100 mm

Sizes, stack counts and substrate sizes from TSMC via Tom's Hardware; the 1.5-reticle generation of 2016 is shown with four stacks, typical of that time. Packaging capacity, not transistor supply, has been one of the tightest bottlenecks of the AI boom.

06 Where to go next

The chip is now a building.

A decade ago a processor was one die in one package. An AI accelerator is a small city of dies, stacked and stitched, and the stitching is as hard as the dies themselves.

Every figure comes from a publisher named below. Package drawings are illustrations, not a specific product.