Presentation
Late Breaking Results: Micro-Dense Tensor-Core-Aligned N:M Sparse Kernel for Efficient Inference Acceleration
DescriptionSemi-structured N:M sparsity theoretically reduces arithmetic cost but often fails to deliver proportional end-to-end latency reduction on modern GPUs. We identify that a key bottleneck lies in irregular execution paths and poor Tensor Core utilization in existing sparse kernels. This work presents MD-SpMM, a Tensor-Core-native N:M sparse CUDA kernel that restructures sparse computation into micro-dense MMA-aligned dataflow. By decoupling sparsity handling from the execution-critical path and introducing inference-scale-aware adaptive parallelism, MD-SpMM restores dense-like execution efficiency while preserving sparsity benefits. Preliminary results across multiple GPU platforms demonstrate up to 2x average speedup over dense cuBLAS and significant gains over state-of-the-art N:M sparse kernels, validating the practicality of Tensor-Core-aligned N:M acceleration.
Event Type
Work in Progress
TimeMonday, July 275:31pm - 5:32pm PDT
LocationExhibit Hall
