Literature index / 15 works

Research literature and open questions

The literature is converging on a more precise question: not whether dense flash can sit near compute, but which workloads survive its latency, access granularity, and write limits.

Open research axes / 01

Open research questions

01

Placement

Which model weights, experts, indexes, or caches should occupy HBF—and for how long?

02

Granularity

Can tensor layout and batching absorb NAND page-level access without wasting bandwidth?

03

Endurance

Which write-heavy inference paths age flash too quickly, and what policies contain that cost?

04

Control

Should data movement be managed by hardware, the runtime, the model server, or all three?

Live debate / 02

Workload suitability and endurance

The strongest 2026 papers do not divide neatly into “for” and “against.” They identify read-dominant residency and large static structures as promising, while treating transient KV caches and small irregular accesses as much harder.

Peer-reviewed / 03

Peer-reviewed literature

Working record / 04

Preprints and workshop papers