Deep Water Semiconductor半导体深水区 · translated column
SemiPulse →

HBF: KV Cache Savior or Trap? Berkeley and FuriosaAI Papers Reveal Scheduling Decides

Two papers published in the same month give opposite verdicts on high-bandwidth flash (HBF). Whether it solves the KV cache bottleneck or creates a trap depends entirely on the scheduler.

October 04, 2026  ·  originally in Chinese

In a data center rack, an inference server’s memory slots are fully occupied, and operators watch queue lengths grow with anxiety: the longer the context and the more agent tasks running, the more KV cache consumes GPU memory. In September, a research team from the University of California, Berkeley, and inference chip company FuriosaAI introduced a new medium into this server’s simulation environment—high-bandwidth flash (HBF).

The conclusion is clear: the exponential growth of KV cache has exceeded the capacity of a single GPU memory tier. A new storage layer is inevitable, but its effectiveness depends on scheduling strategy; the medium itself accounts for only half the equation. The evidence comes from two papers—one giving it high marks, the other a poor review—which together define the boundaries.

Why is a new tier necessary? On the supply side, GPU memory capacity and pricing have lost elasticity in the face of AI demand—the three major manufacturers are shifting production lines toward HBM, driving up general memory prices, and the era of pricing by 'cost per GB' pressures all downstream users. On the demand side, large model context windows are stretching, and agentic workloads have turned 'compute-and-discard' inference into a norm of 'retain-and-reuse': intermediate results, tool outputs, and system prompts from a single task must be saved as KV state for the next call. The demand curve for capacity and the price curve for GPU memory are moving in opposite directions, necessitating a storage layer that is cheaper than GPU memory but significantly faster than standard solid-state drives. HBF targets this gap.

The Case For: Measured Numbers in a Three-Tier Architecture

The paper by Berkeley and FuriosaAI, titled 'Characterizing High Bandwidth Flash for LLM Serving,' was posted to the paper platform in September. First, the medium itself: HBF was proposed by flash memory manufacturers in 2025, with the concept of stacking multi-layer NAND dies and connecting them directly to the accelerator’s memory path via a wide interface within the package, bypassing PCIe and NVMe links—capacity is calculated as flash, while read latency and bandwidth are calculated as near-memory. The research team built a 'GPU memory—HBF—host' three-tier storage system, paired with a cache-aware buffering scheduling strategy, and ran simulations using real workload traces. Results: the fastest HBF-enhanced scheme reduced completion time by 36.1% to 87.0% compared to a pure GPU memory scheme; energy consumption was saved by up to 55.8%, though light-load scenarios actually consumed more power; with scheduling, HBF write endurance estimates extended from 4.77 years to 14.82 years.

The Bear Case: Flash SSD Usage Slows the Entire Stack

Criticism stems from a separate paper released in August by a team from Peking University, which asks directly: Is HBF bad? Treating HBF as a larger SSD substitute for offloading KV cache, they tested five models across four production trajectories. The results showed increased end-to-end latency and reduced effective throughput, with latency more than doubling in the heaviest scenarios. Three factors explain this: the flash capacity gained by HBF crowds out near-memory capacity and bandwidth, where the sacrifice outweighs the medium's benefits; traffic reaching the HBF layer is write-heavy, hitting the weakest point of this medium's write endurance; and 16-layer TLC stacking operates near the 80-degree Celsius junction temperature limit, resulting in actual bandwidth far below the interface peak. The team's conclusion is equally clear: the problem lies not in the medium, but in the usage—HBF should function as a dedicated tier with on-demand selection, reuse awareness, and write budgets.

One data point in the bear case is particularly noteworthy for hardware engineers: under maximized ideal assumptions, moving the attention compute engine into the HBF base die yields almost zero benefit, whereas placing the same engine on the memory base die yields over 40% improvement—because only just over 10% of the KV share actually resides in HBF. The 'move compute to the data' approach hits a wall of data distribution at the flash tier; the base die area budget should be spent on control and scheduling. This draws a red line for a whole class of near-memory compute designs: first determine which tier the data resides in, then decide where to place the compute.

36-87%

The Bull Case: Completion Time Reduction

4.77→14.82 years

Post-scheduling write lifespan

80°C

The Bear Case: Thermal Limit Constraint

The three figures above come from two papers: the completion time reduction of the HBF-enhanced system in the Berkeley proposal, the extent to which scheduling strategies extend HBF write lifespan, and the thermal limit temperature reached by 16-layer TLC stacking in Peking University's real-world tests—the first two support the bull case, while the last explains the bear case.

The Answer from Reading Both Papers Together

Reading the two papers side by side, the divergence actually converges on a single criterion: the reuse structure of KV data. Agentic workloads repeatedly return to the same context; the more concentrated the reuse, the more worthwhile it is to use scheduling to keep hot data in memory and sink cold data to HBF. All the benefits in the bull case come from this 'synergy of data placement and scheduling.' The bear case, by contrast, measures 'indiscriminate offloading,' where cold tail data with heavy write traffic hits the endurance weakness. The same medium: the scheduler determines whether it is an accelerator or a drag.

The author's judgment: by 2028, a packaged flash tier will appear in the memory roadmaps of mainstream AI accelerators—the growth curve of KV cache has already crossed the boundary of memory capacity economics, and this tier has real demand underpinning it. The falsification condition is also clear: if by the end of 2028 no accelerator or memory vendor has written HBF-like media into their mass-production product roadmap, it will mean that scheduling benefits cannot support the engineering costs, and the new-tier thesis will have failed.

Two papers chart distinct paths for players in the storage value chain: memory makers should focus on media refinement and packaging, while system and chip vendors must prioritize scheduling software. For memory manufacturers, the opportunity lies in optimizing the medium and packaging—specifically stacking layers, read bandwidth, and write endurance, which define the usable boundaries of this tier. Extending write endurance from five to under fifteen years creates scheduling headroom that enables TLC particles to enter the market. For system and chip vendors, the opportunity is in scheduling software capabilities such as data placement, cache awareness, and write budget control. These features currently exist only in academic papers and have not yet been implemented as product-grade features. The first vendor to make this a default switch in inference frameworks will hold the power to define this new tier.

Media launches happen every year, and scheduling papers grow thicker annually. The next reshuffling of storage tiers will likely begin in the scheduler code before packaging manufacturers move.

This is an automated English translation of a column originally published in Chinese as《半导体深水区》. Numbers and product names are preserved from the original; wording is machine-generated and may differ from the author's intent. ← All articles