Presentation
Late Breaking Results: Hardware-Efficient PQ Search Framework for Sparse Coding-Based KV Cache Compression
DescriptionThis paper proposes an OPQ-based two-stage search framework for efficient sparse coding in KV cache compression. By integrating Optimized Product Quantization (OPQ) with a filter-and-refine strategy, our framework identifies a compact candidate subset using compressed metadata before performing exact refinement. Experimental results show that the proposed method reduces computational volume by up to 2.46x and memory traffic by 4.5x compared to the Batch-OMP baseline, while maintaining exact support recovery with only 12.5% memory overhead.
Event Type
Work in Progress
TimeMonday, July 276:15pm - 6:16pm PDT
LocationExhibit Hall
