BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Los_Angeles
X-LIC-LOCATION:America/Los_Angeles
BEGIN:DAYLIGHT
TZOFFSETFROM:-0800
TZOFFSETTO:-0700
TZNAME:PDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0700
TZOFFSETTO:-0800
TZNAME:PST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260730T152640Z
LOCATION:Exhibit Hall
DTSTART;TZID=America/Los_Angeles:20260728T174300
DTEND;TZID=America/Los_Angeles:20260728T174300
UID:dac_DAC 2026_sess306_LBR088@linklings.com
SUMMARY:Late Breaking Results: Structured Expert Routing for Efficient MoE
  Inference under GPU–CPU Orchestration
DESCRIPTION:Yixiao Chen, Arman Akbari, and Arash Akbari (Northeastern Univ
 ersity); Zhendong Mi (Stevens Institute of Technology); Enfu Nan (Northeas
 tern University); Xiaowei Lin (ETH Zurich); Weiwei Chen (EmbodyX Inc.); Sh
 aoyi Huang and Hao Wang (Stevens Institute of Technology); and Pu Zhao and
  Yanzhi Wang (Northeastern University)\n\nHybrid GPU–CPU deployment of lar
 ge Mixture-of-Experts (MoE) models is often latency-bound by CPU–side expe
 rt execution and cross-device orchestration overhead. To overcome this lim
 itation, we propose a structured expert routing strategy, implemented via 
 asymmetric expert skipping, which reduces expert activation in latency-cri
 tical layers while applying lightweight magnitude compensation and norm ca
 libration to preserve model accuracy. Experiments on quantized Qwen3 (30B 
 and 235B) models show that our method improves end-to-end generation throu
 ghput by up to 60% while achieving 3.6% higher accuracy than uniform exper
 t reduction under the same compute budget.\n\nTrack: Student\n\n
END:VEVENT
END:VCALENDAR
