Organization <Full Program · Contributors · Organizations · Search Program · Flagged · My Recommendations · Happening NowMore…Search ProgramFlaggedMy RecommendationsHappening NowHarbin Institute of Technology, ShenzhenPresentationsResearch ManuscriptExpertflow: Efficient Mixture-of-Experts Inference via Predictive Expert Caching and Token Scheduling4:23pm - 4:36pm PDT Wednesday, July 29 Mtg Room 201BAIAI5-II. AI/ML System and Platform DesignResearch ManuscriptZipfloat: An Ultra-Fast and Lossless Floating-Point Compressor for LLM Inference Systems4:42pm - 4:54pm PDT Tuesday, July 28 Mtg Room 101BAIAI5-I. AI/ML System and Platform DesignContributorsHao HuShaohuai ShiKaijie TangShihao WangWen XiaXiangyu Zou