Presentation
Late Breaking Results: Towards Low-Latency TinyML via Regularized Activation Packing
DescriptionIn recent years, customized and low-precision CNNs have emerged,
tailored for TinyML applications and well-suited for FPGA deploy-
ment. DSP Packing has been proposed to increase the computa-
tional density of the limited DSP blocks on modern FPGAs, but
current solutions under-utilize key DSP components, like the pre-
adder. In this work, we introduce R-Pack, a novel activation- and
weight-packing technique, able to fully utilize DSP blocks, doubling
their multiplication density to boost inference performance. Our
evaluation showcases our framework's capabilities in reducing the
inference latency by 80% on average, for a mean accuracy loss of
only 2.23%, compared to the baseline state-of-the-art hls4ml tool.
tailored for TinyML applications and well-suited for FPGA deploy-
ment. DSP Packing has been proposed to increase the computa-
tional density of the limited DSP blocks on modern FPGAs, but
current solutions under-utilize key DSP components, like the pre-
adder. In this work, we introduce R-Pack, a novel activation- and
weight-packing technique, able to fully utilize DSP blocks, doubling
their multiplication density to boost inference performance. Our
evaluation showcases our framework's capabilities in reducing the
inference latency by 80% on average, for a mean accuracy loss of
only 2.23%, compared to the baseline state-of-the-art hls4ml tool.
Event Type
Late Breaking Results
TimeMonday, July 275:51pm - 5:55pm PDT
LocationExhibit Hall
Similar Presentations
