Presentation
NM-IF: A Deployable Near-Memory Interface for Incremental System-to-Memory Offload
DescriptionModern computing systems are processor-centric, yet most hardware resources are devoted to storage and movement, which make data transfer a dominant performance bottleneck.
Recently, processing-near-memory (PnM) architectures have emerged in both commercial products and research prototypes.
However, practical PnM adoption is bottlenecked by the lack of a deployable system-to-memory interface that minimizes host involvement while managing offloading and memory access.
To overcome processor-centric adoption barriers, we propose NM-IF, a deployable near-memory interface that enables non-intrusive offloading without requiring system redesign.
NM-IF adopts a data-centric classification of near-memory traffic to streamline the management of read-mostly resident parameters, streaming activations, and on-demand output writebacks with data-aware mapping.
Additionally, it executes parameters and activations in a deterministic phase schedule to enable parallelism, while processing outputs in an event-driven manner to avoid unnecessary periodic scheduling.
Experiments show that NM-IF eliminates the 3×/7× latency growth of 8-bit/4-bit unpacking while maintaining a constant per-fetch mapping cost. It enables effective amortization of the fixed datapath cost for 64-bit system-bus fetches packing beyond 21 elements, improving per-element efficiency.
NM-IF further overlaps parameter/activation handling (22 cycles) with a 1-cycle output writeback, resulting in 23 cycles total.
Overall, NM-IF offers a practical near-memory interface that enhances both performance and efficiency at runtime.
Recently, processing-near-memory (PnM) architectures have emerged in both commercial products and research prototypes.
However, practical PnM adoption is bottlenecked by the lack of a deployable system-to-memory interface that minimizes host involvement while managing offloading and memory access.
To overcome processor-centric adoption barriers, we propose NM-IF, a deployable near-memory interface that enables non-intrusive offloading without requiring system redesign.
NM-IF adopts a data-centric classification of near-memory traffic to streamline the management of read-mostly resident parameters, streaming activations, and on-demand output writebacks with data-aware mapping.
Additionally, it executes parameters and activations in a deterministic phase schedule to enable parallelism, while processing outputs in an event-driven manner to avoid unnecessary periodic scheduling.
Experiments show that NM-IF eliminates the 3×/7× latency growth of 8-bit/4-bit unpacking while maintaining a constant per-fetch mapping cost. It enables effective amortization of the fixed datapath cost for 64-bit system-bus fetches packing beyond 21 elements, improving per-element efficiency.
NM-IF further overlaps parameter/activation handling (22 cycles) with a 1-cycle output writeback, resulting in 23 cycles total.
Overall, NM-IF offers a practical near-memory interface that enhances both performance and efficiency at runtime.
Event Type
Engineering Poster
TimeTuesday, July 285:00pm - 6:00pm PDT
LocationDAC Pavilion, Exhibit Floor
