Close

Presentation

MBA: Mega-Buffer Architecture Based on 1T1C IGZO eDRAM for Efficient LLM Inference
DescriptionSubstantial weight and KV cache transfers between the GPU and HBM severely constrain Large Language Model (LLM) inference system performance. While LLM compression techniques can reduce memory footprint and transfer volume, existing methods are often task-specific and compromise model accuracy. This work proposes a mega-buffer architecture (MBA) that integrates gigabyte-scale on-chip buffers utilizing the embedded DRAM (eDRAM) technology recently introduced by TSMC. To fully leverage this large on-chip memory, we introduce a KV cache prioritized mapping (KVP) scheme that minimizes inefficient KV cache traffic between the HBM and the chip. Furthermore, a highly efficient pipeline integrating double buffering mechanism (DB) is co-designed with an iteration-aware eviction strategy (IA) to enhance data reuse and sustain high compute utilization. Evaluation results show that MBA attains 5.99x and 3.63x end-to-end speedup, and 8.86x and 4.93x energy efficiency improvement against GPU and LLM accelerator baselines, demonstrating a highly efficient architectural solution for LLM inference.