BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Los_Angeles
X-LIC-LOCATION:America/Los_Angeles
BEGIN:DAYLIGHT
TZOFFSETFROM:-0800
TZOFFSETTO:-0700
TZNAME:PDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0700
TZOFFSETTO:-0800
TZNAME:PST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260730T152640Z
LOCATION:Exhibit Hall
DTSTART;TZID=America/Los_Angeles:20260728T174600
DTEND;TZID=America/Los_Angeles:20260728T174700
UID:dac_DAC 2026_sess306_LBR144@linklings.com
SUMMARY:Late Breaking Results: Audio-Language-Action Models with Direct Sp
 eech Conditioning for Robotic Manipulation
DESCRIPTION:Enfu Nan, Pu Zhao, Yixiao Chen, and Lin Zhao (Northeastern Uni
 versity); Juyi Lin (NEU); Chen Wang and Weiwei Chen (EmbodyX Inc.); and Ya
 nzhi Wang (Northeastern University)\n\nVision-language-action (VLA) models
  enable robots to follow instructions in typed text, limiting broader depl
 oyment in natural speech.The traditional solution by prepending a speech r
 ecognition before VLA to convert speech into texts would incur additional 
 latency and propagate errors. To address this, we present Speech-pi_0.5, w
 hich maps raw speech directly to motor commands without explicit textual t
 ranscriptions. Specifically, built on pi_0.5, we use  a cross-modal projec
 tion to convert voice frames directly into audio tokens with significant r
 esponse latency reduction. To further improve the accuracy, we adopt split
  LoRA adaptation with dedicated audio and task adapters, under a two-stage
  training. Speech-pi_0.5 achieves 3\times lower response latency than  ASR
   pipeline while maintaining competitive task success on LIBERO, and gener
 alizes  robustly to unseen accents where a single adapter  degrades signif
 icantly.\n\nTrack: Student\n\n
END:VEVENT
END:VCALENDAR
