Session
Fast, Furious, and Fault-Tolerant: Accelerating the Generative Grind
DescriptionThis session highlights cutting-edge techniques for accelerating modern AI workloads through architectural innovation, sparsity, and model-aware optimization. The papers explore efficient implementations spanning stochastic computing, sparse and quantized inference, diffusion and large language models, as well as GPU-based sparse linear algebra. Together, they demonstrate how hardware–algorithm co-design can unlock substantial performance and efficiency gains for emerging AI applications from edge devices to large-scale systems.
Event Type
Research Manuscript
TimeWednesday, July 291:30pm - 3:00pm PDT
LocationMtg Room 101B
AI
AI3-II. AI/ML Application and Infrastructure
Presentations
