Presentation
3D-ReSAC: A Near-3D-Stacked-DRAM Processor with Area-Efficient Sequential Refresh-Skip and High-Link-Utilization Arrowmesh Communication for Fast Low-Batch LLM Inference
DescriptionLow-batch LLM inference on edge hardware places stringent demands on both memory bandwidth and computational capacity. While 3D-stacked DRAM accelerators offer a promising solution, they introduce two critical overheads that are frequently under-optimized: DRAM refresh and collective communication.
To mitigate these issues, we propose 3D-ReSAC, a near-memory processor based on 3D-stacked DRAM, equipped with an area-efficient sequential refresh-skip method and high-link-utilization ArrowMesh communication.
Our evaluations show that 3D-ReSAC reduces refresh and communication overheads by 7–100% and 53–75%, respectively, leading to a 1.12× to 2.02× latency reduction across low-batch LLM inference workloads.
To mitigate these issues, we propose 3D-ReSAC, a near-memory processor based on 3D-stacked DRAM, equipped with an area-efficient sequential refresh-skip method and high-link-utilization ArrowMesh communication.
Our evaluations show that 3D-ReSAC reduces refresh and communication overheads by 7–100% and 53–75%, respectively, leading to a 1.12× to 2.02× latency reduction across low-batch LLM inference workloads.
Event Type
Research Manuscript
TimeTuesday, July 2810:43am - 10:56am PDT
LocationMtg Room 203AB
