BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Los_Angeles
X-LIC-LOCATION:America/Los_Angeles
BEGIN:DAYLIGHT
TZOFFSETFROM:-0800
TZOFFSETTO:-0700
TZNAME:PDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0700
TZOFFSETTO:-0800
TZNAME:PST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260730T152640Z
LOCATION:Exhibit Hall
DTSTART;TZID=America/Los_Angeles:20260728T175000
DTEND;TZID=America/Los_Angeles:20260728T175000
UID:dac_DAC 2026_sess306_WIP3279@linklings.com
SUMMARY:Nucleus: A Reconfigurable Long-Context LLM Accelerator Using Adapt
 ive Outlier-Aware KV Cache Quantization
DESCRIPTION:Youngmin Cho (Ajou University), Yoontae Lee (University Colleg
 e London), Lojin Park (KAIST (Korea Advanced Institute of Science and Tech
 nology)), Sunjae Lee (Sungkyunkwan University), Jimin Lee (Carnegie Mellon
  University), Jeongwoo Park (Sungkyunkwan University), and Young Oh (Ajou 
 University)\n\nLarge Language Models (LLMs) face significant compute and m
 emory bottlenecks from massive key-value data in long-context inference. W
 e present Nucleus, a configurable accelerator that applies online outlier-
 aware quantization to reduce TTFT and TBT. Its runtime-configurable outlie
 r detector dynamically adapts to KV distribution patterns, enabling precis
 ion-aware computation. With dynamic outlier compaction and a reconfigurabl
 e fused multiply-accumulate (FMA) engine, Nucleus achieves near-ideal scal
 ing across quantization levels. A network-on-chip (NoC) scheduler coordina
 tes multiple Nucleus cores to overlap single-batch prefill and multi-batch
  decode, balancing variadic-length workloads. Fabricated in Samsung 5-nm t
 echnology, Nucleus outperforms state-of-the-art baselines such as NVIDIA A
 100 with KIVI, delivering latency-optimized long-context LLM inference.\n\
 nTrack: Student\n\n
END:VEVENT
END:VCALENDAR
