代码 · 已核验官方资源
审核码:verified_official; 官方关系:verified_mit_training_code_and_logs_without_pretrained_policy; 最近核验:2026-07-28代码:25 tracked files with HumanoidBench, MuJoCo Playground and IsaacLab paths
权重:not shipped
数据:official training logs shipped
许可证:MIT;MIT source with retained notices for adapted upstream code.
依赖:PyTorch、HumanoidBench、MuJoCo Playground fork、IsaacLab 4.5.0、Booster Gym fork
硬件/传感器:Paper uses one NVIDIA A100 and 16 CPU cores、default full-size humanoid setup can require 40-50 GB VRAM、Booster T1 for author physical deployment
实机证据:Author qualitative 12-DoF Booster T1 sim-to-real deployment; no physical trial denominator.
独立复现:mixed simulation reproduction
限制/反面证据:No official pretrained policy;Very high VRAM at default scale;Reward/action tuning is algorithm-specific;Final checkpoint sim-to-sim failure report;No independent physical reproduction
当前建议:高吞吐人形强化学习主流技术2.0·先复算仿真;Use training-log curves as a smoke-test target, store executed clamped actions consistently and validate an earlier checkpoint before the final checkpoint.
重点标签:off-policy-RL、code:verified、logs:verified、community:mixed
核验来源:来源 1 · 来源 2 · 来源 3 · 来源 4