Presentation
Late Breaking Results: Audio-Language-Action Models with Direct Speech Conditioning for Robotic Manipulation
DescriptionVision-language-action (VLA) models enable robots to follow instructions in typed text, limiting broader deployment in natural speech.The traditional solution by prepending a speech recognition before VLA to convert speech into texts would incur additional latency and propagate errors. To address this, we present Speech-pi_0.5, which maps raw speech directly to motor commands without explicit textual transcriptions. Specifically, built on pi_0.5, we use a cross-modal projection to convert voice frames directly into audio tokens with significant response latency reduction. To further improve the accuracy, we adopt split LoRA adaptation with dedicated audio and task adapters, under a two-stage training. Speech-pi_0.5 achieves 3\times lower response latency than ASR pipeline while maintaining competitive task success on LIBERO, and generalizes robustly to unseen accents where a single adapter degrades significantly.
Event Type
Late Breaking Results
TimeMonday, July 275:12pm - 5:16pm PDT
LocationExhibit Hall
