论文记录 · 证据边界公开

Dual-Flow Reinforcement Learning with State-Aware Exploration

arXiv (submitted to IEEE; v1 2026-06-29) · 2026-06-29 · 流匹配 actor-critic(回报分布+多模态策略联合建模)vs diffusion/flow RL

论文元数据

作者
Qijun Li、Zheng Fu、Qi Song、Yifei He、Weitao Zhou、Kun Jiang、Diange Yang
机构
机构尚未补齐
发布日期精度
研究方向
流匹配 actor-critic(回报分布+多模态策略联合建模)vs diffusion/flow RL
机器人
仿真连续控制(DeepMind Control Suite、Humanoid-Bench)
DOI
未披露
arXiv
2606.29820
引用元数据
0
记录状态
included
质量检查
无自动质量红旗