Presentation
3D-DuRA: Accelerating Next-Resolution Generation via 3D Near/in-Memory Architecture with Dual-Ring Sparse Attention
DescriptionVisual Autoregressive (VAR) model, via innovative next-resolution prediction, demonstrates significant potential of GPT-style AR models in image generation. However, due to its coarse-to-fine nature, the input token-map size grows dramatically with each step, resulting in excessive memory access and computational overhead. In this paper, we propose 3D-DuRA, an algorithm-architecture co-design based on a hybrid 3D near-memory and in-memory computing architecture equipped with dual-ring sparse attention, for efficient next-resolution visual-autoregressive generation. Experimental results demonstrate that our proposed 3D-DuRA achieves 4.1× improvement in area efficiency compared with RTX 6000 Ada GPU, along with 3.5× and 9.1× speedups and 10.1× and 13.1× improvements in energy efficiency on Infinity2B and VAR-d36, respectively.
Event Type
Research Manuscript
TimeTuesday, July 2812:16pm - 12:30pm PDT
LocationMtg Room 203AB
Similar Presentations
