Presentation
Procyon: Promoting Fine-Grained Multi-Tenancy to Optimize Sparse Streaming Accelerators
DescriptionSparsity has become a defining characteristic of modern workloads, motivating the development of specialized accelerators to mitigate the challenges inherent to sparsity. Sparse streaming accelerators have proved to be an effective solution, yet current designs are restricted to single-workload execution. Additionally, they exhibit significant underutilization in processing elements (PEs) due to their underlying non-zero scheduling strategies. To address these limitations, we propose Procyon, a fine-grain multi-tenancy framework that fuses the PE instruction streams of multiple workloads into a unified execution schedule. This allows instructions from different workloads to execute on the same PE, thereby, improving its utilization and enabling concurrent execution of multiple workloads on a single accelerator instance. We evaluate Procyon on AMD Alveo U55C FPGA using workloads from the SuiteSparse dataset and show that it substantially reduces PE underutilization that results in 3x speedup over state-of-the-art sparse streaming accelerators (Serpens and Chason), and reaches a peak throughput of 61.2 GFLOP/s.
Event Type
Research Manuscript
TimeMonday, July 273:54pm - 4:06pm PDT
LocationMtg Room 202C
