论文记录 · 证据边界公开

Deployable Human Preference Alignment in Robotics: Learning Representative Rewards from Diverse Human Preferences

arXiv technical report · 2026-07-14 · preference aligned robot policy and offline reinforcement learning

论文元数据

作者
Taehyung Kim、Gwangmo Lee、Minjun Chang、Sunghyun Lim、Jongeun Choi
机构
机构尚未补齐
发布日期精度
研究方向
preference aligned robot policy and offline reinforcement learning
机器人
D4RL MuJoCo Ant、D4RL MuJoCo HalfCheetah、D4RL MuJoCo Hopper、D4RL MuJoCo Walker2d
DOI
未披露
arXiv
2607.12466
引用元数据
0
记录状态
included
质量检查
无自动质量红旗