代码 · hold_missing_license
审核码:hold_missing_license; 官方关系:unverified_for_reproducible_open_use; 最近核验:2026-07-20代码:Public repository page and immutable head were inspected, but the promotion gate was not met.
权重:Not promoted or separately verified in this review.
数据:Not promoted or separately verified in this review.
许可证:not_disclosed;README identifies arXiv 2601.21363 and the exact ICLR 2026 paper, but no root license file or GitHub-detected license was found.
依赖:未披露
硬件/传感器:未披露
实机证据:未披露
独立复现:No qualifying independent reproduction with a commit, disclosed hardware and quantitative metrics was established.
限制/反面证据:README identifies arXiv 2601.21363 and the exact ICLR 2026 paper, but no root license file or GitHub-detected license was found.
当前建议:证据观察;README identifies arXiv 2601.21363 and the exact ICLR 2026 paper, but no root license file or GitHub-detected license was found.
重点标签:promotion:withheld、license:missing-or-anomalous、runtime:not-reproduced
核验来源:来源 1
official_code_and_three_model_files_root_license_missing · hold_code_and_weights_license_missing
审核码:hold_code_and_weights_license_missing; 官方关系:verified_official_repository_code_weights_but_root_license_missing; 最近核验:2026-07-20代码:实质代码已发布。提交 32fab57 的 847 个树条目中包含 381 个 Python 文件,覆盖 MuJoCo Playground SAC 预训练、Brax 评测、世界模型训练、微调、JAX 到 Torch 转换和 T1 MuJoCo 播放。
权重:3 个文件端点已匿名核验:T1 rough-terrain zero-shot policy 27,313,236 字节;T1 sim-finetune policy 25,543,932 字节;T1 physics-informed world model 12,945,831 字节。本批未反序列化或运行。
数据:训练脚本可生成 replay buffer,仓库没有发布论文训练 buffer、真实机器人微调轨迹或独立数据集快照。
许可证:not_disclosed;根目录没有 LICENSE、LICENSE.md、COPYING 或 COPYING.md。mujoco_playground 子目录的 Apache-2.0 文件只覆盖其上游代码,不能推定覆盖 LIFT 新增实现、三个模型文件或机器人资产。
依赖:Ubuntu 22.04 and Python 3.10、JAX 0.4.35 CUDA 12 and PyTorch 2.6.0 CUDA 12.4、MuJoCo Playground, Brax, Optax, TorchRL, Hydra, Optuna and Weights & Biases、BoosterGym for the documented Booster T1 deployment
硬件/传感器:README reports testing on NVIDIA H800 and RTX 4090、project page reports several pretraining tasks solved within one hour on RTX 4090、Booster T1 and Unitree G1 models; physical fine-tuning is demonstrated on Booster T1
实机证据:The author project reports zero-shot Booster T1 deployment on indoor/outdoor surfaces and real-world finetuning from 80 to 590 seconds of collected data.;The repository documents model conversion and points to BoosterGym, but does not include a complete guarded real-robot finetuning controller or hardware safety protocol.
独立复现:GitHub issue #2 reports a detailed independent run on Ubuntu 22.04, Python 3.10.20, RTX A6000, CUDA 12.4: 48M-step pretraining reached reward 96 and Brax zero-shot reward 77, while world-model finetuning hit dtype errors, reward collapse/recovery and a 10k-step stall. The author responded that at least 40k steps are needed and pushed a new commit. This is valuable partial reproduction and failure evidence, not a completed comparable finetuning result.
限制/反面证据:代码和三个模型文件缺少根许可证,不能标为已核验开源或可商用。;真实微调轨迹、replay buffer 和安全控制边界未发布。;社区报告表明检查点恢复时 optimizer dtype 易被错误转换,微调可能发生奖励坍塌或长时间卡住。;依赖栈同时包含 JAX/CUDA、PyTorch/CUDA、MuJoCo Playground 与 Brax,版本和显存成本较高。
当前建议:近期高潜:主流强化学习技术 2.0,开放许可待补;大规模 off-policy 预训练与物理先验世界模型微调的组合符合近期高权重路线,且已有社区部分复现;但许可证、稳定微调和真机安全链缺失,在这些问题关闭前只能列为高潜观察。
重点标签:recency:2026-ICLR、关联:主流强化学习技术2.0、关联:近期热门、关联:社区验证、关联:可能成为新主流、cross-domain:SAC-plus-world-model、artifact:code-and-models-present-license-missing、community:partial-reproduction-with-failure
核验来源:来源 1 · 来源 2 · 来源 3 · 来源 4 · 来源 5 · 来源 6