Fine-Tuning VLAs with Self-Demonstrated Generative Control for Multi-Task Manipulation
Abstract
State-of-the-art VLAs such as π0.5 exhibit strong semantic understanding, instruction following and task behavior. However, when deployed on new robots, even minor mismatches in hardware configuration relative to pretraining causes severe performance drops. Finetuning the VLA on in-domain expert data improves performance on the target task but leads to a loss in the original instruction following capabilities of the VLA, as well as overfits to the expert task behavior. In this paper, we propose a self-supervised method that generates online interaction rollouts from the zero-shot VLA as additional training data for VLA finetuning. Our experiments show this finetuning scheme yields strong multi-task policies that (1) improve expert data performance and sample efficiency, (2) inherit pretraining task behavior distilled from the zero-shot model without expert data, and (3) generalize to out-of-distribution objects. We demonstrate the success of our approach on (1) a real ALOHA robot and (2) a new simulation evaluation protocol in RoboTwin.
Results
Citation
Citation withheld for double-blind review.