BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Los_Angeles
X-LIC-LOCATION:America/Los_Angeles
BEGIN:DAYLIGHT
TZOFFSETFROM:-0800
TZOFFSETTO:-0700
TZNAME:PDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0700
TZOFFSETTO:-0800
TZNAME:PST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260730T152640Z
LOCATION:Exhibit Hall
DTSTART;TZID=America/Los_Angeles:20260728T175800
DTEND;TZID=America/Los_Angeles:20260728T175900
UID:dac_DAC 2026_sess306_WIP3306@linklings.com
SUMMARY:A 1.31 TOPS/W CGRA Hardware with SIMD and Programmable Memory Addr
 ess Generation Unit for Edge AI
DESCRIPTION:Rakshith Harish and Vishnu Nambiar (Institute of Microelectron
 ics, Agency for Science, Technology and Research); Yuntian Liu (School of 
 Electrical and Electronic Engineering, Nanyang Technological University); 
 Yi Sheng Chong (Institute of Microelectronics, Agency for Science, Technol
 ogy and Research); Wang Ling Goh (School of Electrical and Electronic Engi
 neering, Nanyang Technological University); and Rahul Dutta and Anh Tuan D
 o (Institute of Microelectronics, Agency for Science, Technology and Resea
 rch)\n\nCoarse-Grained Reconfigurable Arrays (CGRA) offers a balance betwe
 en a processor's flexibility and a domain-specific accelerator's high ener
 gy efficiency. However, conventional CGRAs typically employ some of its Pr
 ocessing Elements (PEs) to compute the addresses for memory access, result
 ing in lower hardware utilization during workload execution. In this work,
  we propose a CGRA featuring a dedicated Address Generation Unit (AGU) tha
 t directly generates addresses for PEs to fetch and send data, thereby ach
 ieving higher utilization. Our simulation results show that incorporating 
 the AGU improves CGRA utilization by a factor of 2× when executing a 64x8 
 square General Matrix Multiplication (GeMM) workload. The proposed CGRA al
 so supports Single-Instruction-Multiple-Data (SIMD) to accelerate operatio
 ns while boosting energy efficiency up to 3.96×. Implemented in 12nm FinFE
 T technology, our post-layout simulation demonstrates that the proposed de
 sign operates at 1 GHz and achieves an energy efficiency of 1.31TOPS/W, wh
 ich is 3.4× better than the state-of-the-art.\n\nTrack: Student\n\n
END:VEVENT
END:VCALENDAR
