Close

Session

Research Manuscript
:
Beyond GPUs: Next-Generation Architectures for Modern AI Workloads
DescriptionRecent advances in artificial intelligence are increasingly driven by specialized hardware architectures tailored to the computational structure of emerging models. This session presents nine papers addressing architectural techniques for improving the efficiency of modern AI workloads, with particular emphasis on large language models (LLMs), diffusion models, and graph neural networks (GNNs). The contributions span both algorithm–architecture co-design and microarchitectural innovations. Collectively, the papers in this session illustrate the rapid evolution of AI hardware toward workload-specific acceleration, tighter algorithm–hardware co-design, and mechanisms that balance performance, energy efficiency, and deployability across both datacenter and edge environments.
Event Type
Research Manuscript
TimeMonday, July 2710:30am - 12:30pm PDT
LocationMtg Room 101B
Topics
AI
Tracks
AI4-I. AI/ML Architecture Design
Presentations
10:30am - 10:43am PDTFusedot: A Multiplication-Fused Dot Product Accelerator for Efficient LLM Inference
10:43am - 10:56am PDTTDH-GNN: An Efficient Topology-Driven Accelerator for Dynamic Heterogeneous GNN
10:56am - 11:10am PDTMASQ: Accelerating Masked Diffusion via Stage-Wise Multi-Precision Quantization
11:10am - 11:23am PDTTAG: A Topology-Aware Architecture for Configurable and Memory-Efficient GNN Acceleration
11:23am - 11:36am PDTFlorella: Accelerating Graph Neural Networks by Sliding Reduction Convention and Flexible Architecture
11:36am - 11:50am PDTAMBER: A Unified Accelerator for Multi-Precision LLM Inference Exploiting Bit-Level Redundancy and Reconfigurability
11:50am - 12:03pm PDTDSPE: An Energy-Efficient Edge Processor for Deepseek Inference with Merkletree-Based Incremental Pruning, Multi-Stage Boothing Lookup and Dynamic Adaptive Posit Processing
12:03pm - 12:16pm PDTFSGen: Agile Fused and Sparse Accelerator Generator with Accurate Power Model for LLM Applications
12:16pm - 12:30pm PDTDM-MARK: Software and Hardware Co-Design of Watermarking Accelerator for Authorized Diffusion Model Usage on Edge Devices