← Feed Deep Dive Matrix Subscribe

HBF for High-Throughput LLM Serving (UC Berkeley, FuriosaAI)

semiengineering.com 2026-10-03
Entities
Companies:FuriosaAI
Technologies:HBFLLM Serving
Industry Analysis
HBF is not a faster flash—it is a structural reordering of the AI inference memory hierarchy. For three years HBM has been the near-exclusive answer to serving bottlenecks, with SK Hynix, Samsung, and Micron holding a pricing oligopoly. The Berkeley–FuriosaAI work inserts a new bandwidth-capacity equilibrium between HBM and NAND, directly challenging the assumption that inference must scale HBM capacity. The technical ripple is immediate: memory controllers, chiplet interconnects, and die-to-die interface standards all require recalibration. FuriosaAI's WASP architecture is natively suited to tiered storage, meaning an inference-ASIC-plus-HBF stack could emerge as a complete alternative to the NVIDIA-HBM ecosystem. On compliance, advanced memory—especially HBM—now sits under multi-country export controls. If HBF follows a NAND process route, its manufacturing complexity is materially lower, offering downstream customers a de-HBM supply path and a fundamental cost-structure shift. Competitively, SK Hynix will likely accelerate HBM4 to preserve its generational lead; Samsung may weaponize V-NAND bandwidth as a defensive play. NVIDIA's CUDA lock-in holds for training, but the inference-side substitution window is opening. The 12–24-month wildcard: whether HBF spawns a standalone inference-memory category, mirroring how GDDR split from DRAM. If so, the memory market bifurcates into a training-HBM / inference-HBF dual-track, and inference-chip vendors gain unprecedented pricing leverage.
Read Original Article →
Related
This page displays AI-generated summaries and metadata for research purposes. Original content belongs to the respective publishers.