Presentation
Optimizing GPU Clock Power: Exploring Register Array Folding and Its Trade-Offs
DescriptionRegister arrays in GPU design consume significant power due to clock tree complexity and addressing logic. Analysis of a key GPU gaming workload revealed that a group of register arrays account for 10.3% of the total sub-block power, making them a prime target for optimization.
Clock power can be reduced by shrinking entry bit width and/or decreasing array depth or count. However, this can impact performance due to reduced capacity increasing stalls, which can introduce backpressure within a design.
Access pattern analysis showed that consecutive arrays are often accessed together, allowing them to be folded into fewer, wider arrays, preserving capacity and avoiding performance loss. This approach, called access pattern-based array folding, reduces addressing logic and clock power. An example optimization achieved a 2.75% net power reduction, 0.63% reduction in standard cell area, and 13.15% reduction in total ICG count.
The solution has broad applicability across industries, including CPU, AI/ML, networking, and SoC, and provides a systematic, architecture‑agnostic method for global clock-tree simplification. By leveraging access patterns to optimize register array design, this novel approach can reduce power consumption while preserving capacity and performance, making it a valuable technique for various designs across the industry.
Clock power can be reduced by shrinking entry bit width and/or decreasing array depth or count. However, this can impact performance due to reduced capacity increasing stalls, which can introduce backpressure within a design.
Access pattern analysis showed that consecutive arrays are often accessed together, allowing them to be folded into fewer, wider arrays, preserving capacity and avoiding performance loss. This approach, called access pattern-based array folding, reduces addressing logic and clock power. An example optimization achieved a 2.75% net power reduction, 0.63% reduction in standard cell area, and 13.15% reduction in total ICG count.
The solution has broad applicability across industries, including CPU, AI/ML, networking, and SoC, and provides a systematic, architecture‑agnostic method for global clock-tree simplification. By leveraging access patterns to optimize register array design, this novel approach can reduce power consumption while preserving capacity and performance, making it a valuable technique for various designs across the industry.
Event Type
Engineering Poster
TimeWednesday, July 293:00pm - 4:00pm PDT
LocationDAC Pavilion, Exhibit Floor
