Presentation
Nucleus: A Reconfigurable Long-Context LLM Accelerator Using Adaptive Outlier-Aware KV Cache Quantization
DescriptionLarge Language Models (LLMs) face significant compute and memory bottlenecks from massive key-value data in long-context inference. We present Nucleus, a configurable accelerator that applies online outlier-aware quantization to reduce TTFT and TBT. Its runtime-configurable outlier detector dynamically adapts to KV distribution patterns, enabling precision-aware computation. With dynamic outlier compaction and a reconfigurable fused multiply-accumulate (FMA) engine, Nucleus achieves near-ideal scaling across quantization levels. A network-on-chip (NoC) scheduler coordinates multiple Nucleus cores to overlap single-batch prefill and multi-batch decode, balancing variadic-length workloads. Fabricated in Samsung 5-nm technology, Nucleus outperforms state-of-the-art baselines such as NVIDIA A100 with KIVI, delivering latency-optimized long-context LLM inference.
Event Type
Work in Progress
TimeMonday, July 276:36pm - 6:37pm PDT
LocationExhibit Hall
