BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Los_Angeles
X-LIC-LOCATION:America/Los_Angeles
BEGIN:DAYLIGHT
TZOFFSETFROM:-0800
TZOFFSETTO:-0700
TZNAME:PDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0700
TZOFFSETTO:-0800
TZNAME:PST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260730T152639Z
LOCATION:Exhibit Hall
DTSTART;TZID=America/Los_Angeles:20260728T170100
DTEND;TZID=America/Los_Angeles:20260728T170200
UID:dac_DAC 2026_sess306_WIP3193@linklings.com
SUMMARY:Late Breaking Results: Micro-Dense Tensor-Core-Aligned N:M Sparse 
 Kernel for Efficient Inference Acceleration
DESCRIPTION:Liu Yiming (University of Science and Technology of China), We
 nqi Lou (USTC), Ke Zhiwei (University of Science&Technology of China), Fen
 grui Zuo and Chao Wang (University of Science and Technology of China), an
 d Xuehai Zhou (USTC)\n\nSemi-structured N:M sparsity theoretically reduces
  arithmetic cost but often fails to deliver proportional end-to-end latenc
 y reduction on modern GPUs. We identify that a key bottleneck lies in irre
 gular execution paths and poor Tensor Core utilization in existing sparse 
 kernels. This work presents MD-SpMM, a Tensor-Core-native N:M sparse CUDA 
 kernel that restructures sparse computation into micro-dense MMA-aligned d
 ataflow. By decoupling sparsity handling from the execution-critical path 
 and introducing inference-scale-aware adaptive parallelism, MD-SpMM restor
 es dense-like execution efficiency while preserving sparsity benefits. Pre
 liminary results across multiple GPU platforms demonstrate up to 2x averag
 e speedup over dense cuBLAS and significant gains over state-of-the-art N:
 M sparse kernels, validating the practicality of Tensor-Core-aligned N:M a
 cceleration.\n\nTrack: Student\n\n
END:VEVENT
END:VCALENDAR
