论文部分内容阅读
基于行为的自主移动机器人在获取外界信息时不可避免地会引入噪声,给其系统性能造成一定的影响。提出了一种基于过程奖赏和优先扫除(PS-process)的强化学习算法作为噪声消解策略。针对典型的觅食任务,以计算机仿真为手段。并与其它四种算法——基于结果奖赏和优先扫除(PS-result)、基于过程奖赏和Q学习(Q-process)、基于结果奖赏和Q学习(Q-result)和基于手工编程策略(Hand)进行比较。研究结果表明比起其它四种算法,本文所提出的基于过程奖赏和优先扫除的强化学习算法能有效降低噪声的影响,提高了系统整体性能。
Behavior-based autonomous mobile robots inevitably introduce noise when obtaining external information, which may affect their system performance to a certain extent. A reinforcement learning algorithm based on process reward and priority sweep (PS-process) is proposed as noise cancellation strategy. For the typical foraging tasks, computer simulation as a means. Based on the results reward and Q-process, result-based reward and Q-result and Hand-based programming strategy )Compare. The results show that compared with the other four algorithms, the proposed reinforcement learning algorithm based on process reward and priority sweep can effectively reduce the impact of noise and improve the overall system performance.