Presentation
Expcheck: Dynamic Expert-Aware Checkpointing for Mixture-of-Experts Based Models
DescriptionLarge-scale Mixture-of-Experts (MoE) models are pivotal in modern AI, yet their massive parameter size creates a "storage wall" for fault tolerance, where limited bandwidth restricts checkpoint frequency and risks significant wasted computation. We present "EXPCheck", a dynamic expert-aware checkpointing system designed to resolve the conflict between massive MoE states and limited persistence bandwidth. Grounded in the observation that expert activation is highly imbalanced, EXPCheck employs a novel "Aging-then-Greedy Expert Selection (AGES)" policy. AGES first enforces an age-based refresh for overdue "cold" experts to prevent indefinite staleness, and then greedily allocates the remaining persistence budget to frequently updated "hot" experts. Implemented on a production-scale training stack, EXPCheck significantly reduces persistence traffic and increases checkpoint frequency by at most 5× compared to full checkpointing, while maintaining downstream model accuracy comparable to standard methods.
Event Type
Research Manuscript
TimeWednesday, July 294:30pm - 4:42pm PDT
LocationMtg Room 203C
Similar Presentations
