该论文记录已去重归并

旧标识 oa-w4416750979 已合并到唯一正式记录。

前往 VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model