Presentation
Designing Power- and Area-Efficient Custom NPUs for Edge AI
DescriptionEdge AI is increasingly popular in applications requiring real-time decision making and autonomous operation. Different from NPUs for cloud platforms, edge AI processors can be made application-specific. By tuning their ISA and memory architecture to the network models required by the application, power consumption and silicon area are drastically reduced.
Tools for application-specific instruction-set processors (ASIPs) can be used to design custom NPUs for edge AI. We present the design of "SmarT", an ASIP with a RISC-V ISA augmented with specialized vector units for convolutions and quantization, with 64 MACs. It supports circular gather/scatter addressing of vector data in parallel with computations. Low-overhead DMA moves data blocks from external to local memory.
ASIP tools enable a software path from TensorFlow using LiteRT. We optimized selected LiteRT kernels in conjunction with the processor architecture. SmarT uses only 200Kgates, while delivering 100GMAC/s performance, making it suited for many low-power sensor, audio and video applications.
Tools for application-specific instruction-set processors (ASIPs) can be used to design custom NPUs for edge AI. We present the design of "SmarT", an ASIP with a RISC-V ISA augmented with specialized vector units for convolutions and quantization, with 64 MACs. It supports circular gather/scatter addressing of vector data in parallel with computations. Low-overhead DMA moves data blocks from external to local memory.
ASIP tools enable a software path from TensorFlow using LiteRT. We optimized selected LiteRT kernels in conjunction with the processor architecture. SmarT uses only 200Kgates, while delivering 100GMAC/s performance, making it suited for many low-power sensor, audio and video applications.
Event Type
Engineering Presentation
TimeWednesday, July 2910:45am - 11:00am PDT
LocationSeaside Ballroom B
