Presentation
Rapid Design and Deployment of High-Radix Switch Fabric for Scale-up and Scale-out AI Systems
DescriptionNext-generation scale-up and scale-out AI systems require high-radix switches that deliver high bandwidth, low latency, and non-blocking performance while scaling to hundreds of ports. Traditional switch fabrics are typically implemented using large crossbars, which satisfy performance requirements but present significant implementation and physical design (PD) challenges. As port counts increase, crossbars become increasingly difficult to develop, floorplan, route, and time, and pose fundamental barriers to scaling across chiplets for continued growth in number of ports.
This work presents NeuraScale, a scalable, non-blocking switch fabric architecture designed to address these challenges. The fabric employs a hybrid topology that combines the scalability and PD regularity of mesh-based structures with Clos-style connectivity to provide system-level non-blocking behavior. The architecture is constructed from repeatable, PD-friendly tiles that can be replicated to build large fabrics, including across chiplets, enabling rapid design iteration and deployment.
Performance evaluation of a 128×800G (102.4 Tb/s) AI switch under permutation traffic demonstrates 100% peak throughput with flat latency, while a conventional mesh saturates at 73% throughput with sharply increasing latency. This approach enables practical realization of high-radix, chiplet-ready AI switches for large-scale AI systems.
This work presents NeuraScale, a scalable, non-blocking switch fabric architecture designed to address these challenges. The fabric employs a hybrid topology that combines the scalability and PD regularity of mesh-based structures with Clos-style connectivity to provide system-level non-blocking behavior. The architecture is constructed from repeatable, PD-friendly tiles that can be replicated to build large fabrics, including across chiplets, enabling rapid design iteration and deployment.
Performance evaluation of a 128×800G (102.4 Tb/s) AI switch under permutation traffic demonstrates 100% peak throughput with flat latency, while a conventional mesh saturates at 73% throughput with sharply increasing latency. This approach enables practical realization of high-radix, chiplet-ready AI switches for large-scale AI systems.
Event Type
Engineering Poster
TimeWednesday, July 293:00pm - 4:00pm PDT
LocationDAC Pavilion, Exhibit Floor
