论文元数据 作者Taehyung Kim、Gwangmo Lee、Minjun Chang、Sunghyun Lim、Jongeun Choi机构机构尚未补齐发布日期精度日研究方向preference aligned robot policy and offline reinforcement learning机器人D4RL MuJoCo Ant、D4RL MuJoCo HalfCheetah、D4RL MuJoCo Hopper、D4RL MuJoCo Walker2dDOI未披露arXiv2607.12466引用元数据0记录状态included质量检查无自动质量红旗