← Back to library
HBF researchPreprintResearch

Catalog record / 14 Aug 2026

DASH gives HBF two concurrent paths into the GPU

DASH combines a direct GPU–HBF route with an HBF–HBM–GPU relay and uses both concurrently to deliver MoE expert weights, while managing immutable experts separately from mutable KV-cache traffic.

Why it matters

It moves beyond treating HBF as passive capacity and directly studies expert placement, early reads, and dual UCIe data paths. The reported 1.94× throughput and 1.90× end-to-end speedup are simulator results from a preprint, not silicon measurements.

Exact source title

Beyond Capacity: Scalable MoE LLM Inference via High-Bandwidth Flash with Direct GPU and HBM Paths

Authors

Seeyeon Kim, Juhyeong Jin, Joo-Young Kim

Tags

DASHMoEexpert weightsdual pathUCIeKV cache

Reading rule

This record preserves the source's maturity label. A roadmap, research result, or announced specification should not be treated as independent evidence of shipping hardware.