📌 Inference Demand Expansion May Diversify the HBM Buyer Structure
New demand for AI servers may be shifting from the training side to the inference side. Dell’s view is that inference demand has already surpassed training, becoming the industry’s pure incremental demand; it expects inference token volume to grow 87-fold by 2030. If enterprises deploy more agent applications, the inference compute required for individual tasks will increase, and so will the consumption of server resources.
The significance for high-bandwidth memory (HBM) lies not only in the potential increase in total volume, but also in the possibility that buyers may become more dispersed. Related forecasts suggest that Nvidia’s share of HBM demand will fall from 58% in 2026 to 41% in 2027, while Google will rise to 31% and AMD to 12% over the same period. If these changes materialize, they would be more consistent with platforms such as TPUs and AMD expanding deployments and jointly absorbing incremental demand, rather than with HBM demand itself weakening.
For the supply chain, the key variable is whether multi-platform inference deployments can translate into actual server deliveries and HBM procurement. A decline in Nvidia’s share does not necessarily mean its demand will decrease, because total demand may still expand; conversely, platform share forecasts cannot be directly equated with revenue growth without support from order, shipment, and memory manufacturer delivery data. Existing information consists mainly of secondary retellings and forecasts and remains to be verified.
▌ Sources