Presentation
HP-CIM: A Computing-in-Memory Transformer Accelerator with ReRAM-Based Hash Predictor for Attention Sparsity Exploitation
DescriptionTransformers excel at sequential modeling but attention and large matrix computation incur high latency and energy. The potential of the hybrid ReRAM-SRAM computing-in-memory(CIM) accelerator is constrained by the redundant attention sparsity. Exploiting sparsity introduces significant overhead in Top-K query-key identification or causes accuracy degradation in prior works. We present HP-CIM, a hybrid accelerator with a ReRAM-based hash predictor(ReHP) that exploits device variability for low-cost projections and couples ReRAM CAM with a K-winner-take-all(K-WTA) circuit for Top-K selection. Furthermore, an optimizable bias-softmax mechanism compensates information loss. Across diverse tasks, HP-CIM delivers 9.05–310.04× energy efficiency and 2.48-16.93× speedups over state-of-the-art CIM-based transformer accelerators.
Event Type
Research Manuscript
TimeWednesday, July 292:47pm - 3:00pm PDT
LocationMtg Room 203AB
Similar Presentations
