Nvidia capacity-allocation loop: from the HBM and CoWoS supply ladder through the rationing lever into hyperscaler fleets, and the reinvestment loop that deepens the lockup
Nvidia rations a supply ladder that is sold out at every rung, converting inelastic AI demand into near-peak margins and free cash flow that is reinvested into deeper packaging and memory lockup. The arrangement is a closed self-reinforcing loop that only breaks when packaging or memory loosens, a rival compute source scales, or hyperscaler capex is cut.
wafer starts
foundry output
sold
out
CoWoS slots
packaging
sold
out
HBM output
allocations
sold
out
2 · every rung sold out
Nvidia
allocation
rationing lever
pre-commits largest
packaging share
GPU + HBM
one NVLink bundle
NVLink spine
sold as one unit
GPU fleet
hyperscaler
AI fleets
1 · hyperscaler AI capex
multi-year commitments
4 · pricing power
near-peak margins
free cash flow
large conversion
5 · reinvest in lockup
packaging + memory commitments
inelastic demand
at the margin
self-reinforcing loop
3
lever engaged · loop intact
CoWoS / HBM loosens
rival compute scales
capex cut
The allocation loop circulates on its own. Flip any unwind condition to open the circuit: the reinvestment feed breaks, the rationed-supply arrow loses power, or the capex commitment is cut. Restore all three and the lever re-engages.