Per-Group Error, Not Total MSE: Fine-Tuning Vision-Language-Action Models for 11-DoF Mobile Manipulation

返回详情
VLA / Vision-Language-Action 每日论文卡