Presentation
Co-Designed Compute Memories for Ultra-Efficient DNN Inference
DescriptionArtificial Intelligence (AI) at the edge must address conflicting performance and efficiency needs. One major challenge is the cost of data movement between processors and memories. Our proposed Near Memory Computing (NMC) architecture tackles this by providing memory arrays with lightweight arithmetic units, tailored to AI algorithms' needs for near sensor processing. We optimize holistically a) the design of NMC hardware, b) its integration within Systems-on-Chip (SoC) components, and c) edge AI applications. Our solution uses data parallelism, inherent in Deep Neural Network (DNN) models, to distribute computations across compute memory banks. It also employs aggressive quantization to reduce weight data size and lower computational demands. Additionally, it integrates seamlessly into standard SoC architectures, enabling end-to-end DNN inference near memory. Performance improvements of up to 250x are achieved compared to software execution, with only about 11% area overhead relative to similar non-compute memories.
Event Type
Research Special Session
TimeTuesday, July 284:00pm - 4:30pm PDT
LocationMtg Room 201A
