Presentation
GREEN: Towards Scalable Energy-Efficient Workload Scheduling and Placement in the Cloud
DescriptionThe workload scheduling and placement problem has been the core of resource management in various distributed systems, including modern cloud computing systems. As modern cloud systems become larger, more heterogeneous, dynamic, complicated, and energy-hungry, developing a scalable, flexible, adaptive, and effective energy-efficient resource manager for such a cloud system is becoming very challenging. Specifically, existing machine learning-based resource managers cannot scale well, and existing heuristic algorithm-based scalable resource managers do not generate the most energy-efficient solutions. To address the limitations in existing methods and solve the problem more effectively, we propose GREEN, a graph neural network-reinforcement learning (GNN-RL) based cloud resource manager that generates high energy-efficiency workload scheduling and placement solutions and scales up to large cloud systems with hundreds and thousands of servers. GREEN reduces this challenging problem to a graph optimization problem and then uses our novel RL formulation and GNN architecture for generating scheduling and placement solutions for cloud systems in various scales. Through extensive cloud simulations using COSCO and real-world experiments using CloudLab, we found that GREEN's solutions save energy by up to 2.17x than those generated by the best previous state-of-the-art (SOTA) resource managers without compromising Service Level Objective (SLO) metrics. Most importantly, GREEN can generate similarly high quality scheduling and placement solutions on systems with 100 to 1000 servers in COSCO cloud simulations.
Event Type
Research Manuscript
TimeTuesday, July 2811:23am - 11:36am PDT
LocationMtg Room 101B
