Close

Presentation

EdgeQ‑GEMM: INT4/INT8 Mixed‑Precision GEMM Accelerator with On‑the‑Fly Quantization for Accuracy–Energy Tunability in Edge AI Processors
DescriptionEdge workloads such as keyword spotting and activity recognition must process continuous FP32 sensor data under tight power–performance–area budgets, making full-precision GEMM units impractical. EdgeQ-GEMM is a compact, processor-integrated INT4/INT8 mixed-precision GEMM accelerator that performs on-the-fly quantization and computation without FP hardware. It supports two modes: adaptive mode, dynamically quantizing FP32 activations to INT4 or INT8 using a design-time QCM to expose a tunable accuracy–energy design space, and layer-wise mode, executing QAT/PTQ models with layer-mixed INT4/INT8 weights by applying on-the-fly INT32-to-INT8 activation quantization. Implemented in edge AI processors and validated via FPGA and 45nm synthesis, EdgeQ-GEMM enables efficient, flexible precision adaptation for diverse workloads.