论文记录 · 证据边界公开

Beyond Imitation: Self-Improving Robot Policies via Off-Policy Q-Planning

arXiv v1 (cs.LG/cs.RO) · 2026-08-21 · 通用或轻量 VLA

论文元数据

作者
Varun Giridhar、Anant Khandelwal、Jeremy A. Collins、Ignat Georgiev、Animesh Garg
机构
affiliations not independently normalized
发布日期精度
研究方向
通用或轻量 VLA
机器人
FastWAM policy、two 6-DoF YAM arms
DOI
10.48550/arXiv.2608.21204
arXiv
2608.21204
引用元数据
0
记录状态
included
质量检查
无自动质量红旗