MuseVLA Multisensor Real-Robot Dataset
MuseVLA 的公开真机多传感器操作数据,固定 Hugging Face revision 提供 RGB、深度、热成像、声学、毫米波雷达、分割掩码、末端 6D 轨迹、手臂/手状态与命令。数据卡为 1,397 条 episode、1,010 条有效、1,930 个标注片段和 11 条指令;论文核心真机训练集为其中 720 条遥操作示范。它是近期多传感器 VLA 与灵巧手数据的可下载证据,但仅研究用途,不是人形全身、视觉触觉、力控或跨本体基准。
- 开放状态
- 已核验开放
- 许可证
- CC-BY-NC-4.0 declared by the official Hugging Face dataset card
- 商业使用
- research use only; the non-commercial dataset license is independent from MuseVLA code/model MIT licenses and from robot/sensor/driver terms
- 人工评审结论
- fixed_revision_public_nongated_real_multisensor_vla_dataset_noncommercial
- 机器人
- table-top robot with Robotera XHand 12-DoF dexterous hand、Intel RealSense RGB-D、two infiRay T2S thermal cameras、Calterah 4T4R 60GHz radar、Sipeed 6+1Mic microphone array
- 任务
- thermal-guided drink pick-and-place, acoustic object search and mmWave radar-guided hidden-object retrieval across towel, clothes, box, item and drink instructions
- 模态
- RGB MP4、depth images、SAM3 segmentation masks、thermal left/right and projected heatmaps、acoustic spectrogram/heatmaps、mmWave radar heatmaps、6D end-effector trajectories、arm/hand command and status、language instruction、segment start/end labels、validity labels
- 格式
- per-episode MP4, JPG/PNG, NPY and TXT streams with JSON labels, top-level instruction/valid-sample indices and normalization statistics
- 采集/生成方式
- MANUS5 glove teleoperation with synchronized multi-sensor recording; synthetic sensor data used in model pretraining is a separate derivative corpus
- 已知问题
- CC-BY-NC-4.0 blocks commercial use; no full local download/checksum, privacy or consent audit, complete arm hardware specification, timing error report, failure ledger, cross-embodiment data or independent physical reproduction. Thermal/audio/radar are not tactile or force channels.
官方真机数据固定版本 · 官方模型固定版本 · 官方代码固定提交 · 论文原文