BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Los_Angeles
X-LIC-LOCATION:America/Los_Angeles
BEGIN:DAYLIGHT
TZOFFSETFROM:-0800
TZOFFSETTO:-0700
TZNAME:PDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0700
TZOFFSETTO:-0800
TZNAME:PST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260730T152640Z
LOCATION:Exhibit Hall
DTSTART;TZID=America/Los_Angeles:20260728T175100
DTEND;TZID=America/Los_Angeles:20260728T175100
UID:dac_DAC 2026_sess306_WIP3281@linklings.com
SUMMARY:PHAP: A Pre-analyzed Head-wise Attention Pattern for Efficient Spa
 rse Attention in LLM Inference
DESCRIPTION:Seonghan Kwon, Daeyoung Kim, Hyunjae Jang, Jaewook Kim, YeonJo
 o Jeong, and Inho Kim (Korea Institute of Science and Technology); Jong-Ko
 ok Kim (Korea University); Seongsik Park (Korea Institute of Science and T
 echnology); and Jongkil Park (KIST)\n\nRecent advancements in large langua
 ge models (LLMs) have demonstrated powerful capabilities across various ap
 plication domains. At the same time, their quadratic computational complex
 ity with respect to input sequence length remains a major bottleneck for e
 fficient inference and deployment. To mitigate this problem, sparse attent
 ion methods utilizing the sparsity of computationally intensive operations
  in attention of LLM were proposed, but still incur additional overhead to
  locate important tokens during inference or performance degradation due t
 o utilizing a fixed pattern, neglecting the difference of head-wise sparsi
 ty. In this paper, we propose a pre-analyzed head-wise attention pattern (
 PHAP) sparse attention method, which constructs head-specific patterns by 
 analyzing position-dependent characteristics and the inherent sparsity var
 iations across attention heads. This one-time pre-analyzed pattern constru
 ction eliminates runtime overhead during inference. In addition, the propo
 sed method constructs optimal static sparsity patterns tailored to the uni
 que characteristics of each attention head, thereby maintaining the baseli
 ne performance. Experimental results show reduced attention computation by
  more than 75% without additional overhead while maintaining performance c
 omparable to full attention in Llama-3. Moreover, it achieves 1.17x faster
  inference than conventional sparse attention methods, demonstrating its e
 ffectiveness in maximizing the computational efficiency of LLMs.\n\nTrack:
  Student\n\n
END:VEVENT
END:VCALENDAR
