Hierarchical state-abstracted and socially augmented Q-Learning for reducing complexity in agent-bas

来源 :Journal of Control Theory and Applications | 被引量 : 0次 | 上传用户:ahjockey
下载到本地 , 更方便阅读
声明 : 本文档内容版权归属内容提供方 , 如果您对本文有版权争议 , 可与客服联系进行内容授权或下架
论文部分内容阅读
A primary challenge of agent-based policy learning in complex and uncertain environments is escalating computational complexity with the size of the task space(action choices and world states) and the number of agents.Nonetheless,there is ample evidence in the natural world that high-functioning social mammals learn to solve complex problems with ease,both individually and cooperatively.This ability to solve computationally intractable problems stems from both brain circuits for hierarchical representation of state and action spaces and learned policies as well as constraints imposed by social cognition.Using biologically derived mechanisms for state representation and mammalian social intelligence,we constrain state-action choices in reinforcement learning in order to improve learning efficiency.Analysis results bound the reduction in computational complexity due to stateion,hierarchical representation,and socially constrained action selection in agent-based learning problems that can be described as variants of Markov decision processes.Investigation of two task domains,single-robot herding and multirobot foraging,shows that theoretical bounds hold and that acceptable policies emerge,which reduce task completion time,computational cost,and/or memory resources compared to learning without hierarchical representations and with no social knowledge. A primary challenge of agent-based policy learning in complex and uncertainty environments is escalating computational complexity with the size of the task space (action choices and world states) and the number of agents. Nonetheless, there is ample evidence in the natural world that high -functioning social mammals learn to solve complex problems with ease, both individually and cooperatively. This ability to solve computationally intractable problems stems from both brain circuits for hierarchical representation of state and action spaces and learned policies as well as constraints imposed by social cognition.Using biologically derived mechanisms for state representation and mammalian social intelligence, we constrain state-action choices in reinforcement learning in order to improve learning efficiency. Analysis result bound the reduction in computational complexity due to state of, hierarchical representation, and socially constrained action selection in agent- based learning problems that can be described as variants of Markov decision processes. Investigation of two task domains, single-robot herding and multirobot foraging, shows that theoretical bounds hold and that acceptable policies emerge, which reduce task completion time, computational cost, and / or memory resources compared to learning without hierarchical representations and with no social knowledge.
其他文献
“不准哭!” “不许看电视!” “不可以吃糖!” “不能这么做!” “不要总是发脾气”……   前段时间,我们在身边的妈妈群里问大家,平时最常和孩子说的话,“不XXX”的句式,上了第一名。这个频次,可不是我们的“感觉”而已,有个研究结果是这么说的:“学步期孩子平均每 9 分钟,就要听到一次‘不’或类似的否定词汇。”   也就是说,差不多每过 10 分钟,孩子就要遭到一次否定。   没有妈妈想要否定
期刊
市直机关各级党组织和广大共产党员积极响应中央和市委的号召,认真落实党的十七大精神和中央关于奥运筹办工作的一系列指示精神,按照科学发展观的要求,充分发挥党组织的战斗
提起酒店洗衣房经理、工会女工委员石宝春,北京金融街威斯汀大酒店的员工们都亲切地称她“石大姐”。今年52岁的石宝春在洗衣房已经工作8年多,不论酷暑严寒,不论烈日风雨,一
从本质上来说,电视新闻主播不仅代表的是自己,更是一个电视台以及地区的形象.为了能够切实提高新闻主播的外在形象,给观众留下深刻的印象,文章主要围绕新闻主播定妆造型方面
请下载后查看,本文暂不支持在线获取查看简介。 Please download to view, this article does not support online access to view profile.
本文报道1例以“低血糖、呼吸不规则”起病的能量代谢异常线粒体病患儿,通过肌肉活检、线粒体基因检测明确诊断。线粒体病临床表现多样,不易诊断,实验室检查特别是病理学和基因
“妈妈,我刚才已经洗手了.”“明明就没洗,你当我眼瞎?”这是大部分父母的第一反应.暴跳如雷、指责、惩罚……rn有一个研究发现,孩子两岁时就会说谎;4-6岁时,九成会说谎;而6
期刊
纳米活性碳纤维(activated carbon nanofiber, ACNF)因其所具有的纳米结构与特殊的孔结构性能被公认为是高性能纤维材料之一,可用于催化剂及催化剂载体、储氢设备、超级电容器
新媒体的不断发展,互联网技术越来越成熟,促使新旧媒体融合趋势越来越明显,在这样的媒体环境下,单一的图片和文字信息已经无法满足受众的基本需求.麦克卢汉说: “媒介是人的
我们现在的日子越来越好,物质条件极为丰富,可是孩子的肥胖率更高、体质更弱,孩子的身体也容易出现各种症状,情绪的性格的问题也多了,为什么会这样?为什么好日子却养出了弱孩
期刊