Presentation
Late Breaking Results: Structured Expert Routing for Efficient MoE Inference Under GPU–CPU Orchestration
DescriptionHybrid GPU–CPU deployment of large Mixture-of-Experts (MoE) models is often latency-bound by CPU–side expert execution and cross-device orchestration overhead. To overcome this limitation, we propose a structured expert routing strategy, implemented via asymmetric expert skipping, which reduces expert activation in latency-critical layers while applying lightweight magnitude compensation and norm calibration to preserve model accuracy. Experiments on quantized Qwen3 (30B and 235B) models show that our method improves end-to-end generation throughput by up to 60% while achieving 3.6% higher accuracy than uniform expert reduction under the same compute budget.
Event Type
Late Breaking Results
TimeMonday, July 275:55pm - 5:58pm PDT
LocationExhibit Hall
Similar Presentations
