Close

Presentation

PRISM: Priority Aware Shared Scale Microscaling Format for 4-Bit Quantization
DescriptionMicroscaling formats have emerged as prominent candidates for 4-bit quantization of modern AI models due to its fine-grained group-wise quantization granularity. However, such format still exhibit fundamental accuracy degradations. In particular, MXFP4 and NVFP4 are limited by fixed shared-scale precision, and NVFP4 further cannot cover the full value range without an added FP32 scale. This paper presents PRISM, a microscaling format with a single 8-bit encoded group level shared scale that adaptively allocates shared scale's reprsentation based on the relative importance of values. Through our evaluation, PRISM surpasses conventional Microscaling format accuracy with only 0.62% area and 0.86% energy overhead.