复杂领域中扩展多智能体强化学习

2008 IEEE/WIC/ACM International Conference on Web Intelligence and Intelligent Agent Technology Pub Date : 2008-12-09 DOI:10.1109/WIIAT.2008.259

D. Xiao, A. Tan

{"title":"复杂领域中扩展多智能体强化学习","authors":"D. Xiao, A. Tan","doi":"10.1109/WIIAT.2008.259","DOIUrl":null,"url":null,"abstract":"TD-FALCON (temporal difference-fusion architecture for learning, cognition, and navigation) is a class of self-organizing neural networks that incorporates temporal difference (TD) methods for real-time reinforcement learning. In this paper, we present two strategies, i.e. policy sharing and neighboring-agent mechanism, to further improve the learning efficiency of TD-FALCON in complex multi-agent domains. Through experiments on a traffic control problem domain and the herding task, we demonstrate that those strategies enable TD-FALCON to remain functional and adaptable in complex multi-agent domains.","PeriodicalId":393772,"journal":{"name":"2008 IEEE/WIC/ACM International Conference on Web Intelligence and Intelligent Agent Technology","volume":null,"pages":null},"PeriodicalIF":0.0000,"publicationDate":"2008-12-09","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"8","resultStr":"{\"title\":\"Scaling Up Multi-agent Reinforcement Learning in Complex Domains\",\"authors\":\"D. Xiao, A. Tan\",\"doi\":\"10.1109/WIIAT.2008.259\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"TD-FALCON (temporal difference-fusion architecture for learning, cognition, and navigation) is a class of self-organizing neural networks that incorporates temporal difference (TD) methods for real-time reinforcement learning. In this paper, we present two strategies, i.e. policy sharing and neighboring-agent mechanism, to further improve the learning efficiency of TD-FALCON in complex multi-agent domains. Through experiments on a traffic control problem domain and the herding task, we demonstrate that those strategies enable TD-FALCON to remain functional and adaptable in complex multi-agent domains.\",\"PeriodicalId\":393772,\"journal\":{\"name\":\"2008 IEEE/WIC/ACM International Conference on Web Intelligence and Intelligent Agent Technology\",\"volume\":null,\"pages\":null},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2008-12-09\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"8\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"2008 IEEE/WIC/ACM International Conference on Web Intelligence and Intelligent Agent Technology\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1109/WIIAT.2008.259\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"2008 IEEE/WIC/ACM International Conference on Web Intelligence and Intelligent Agent Technology","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/WIIAT.2008.259","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 8

摘要

TD- falcon(用于学习、认知和导航的时间差异融合架构)是一类自组织神经网络，它结合了用于实时强化学习的时间差异(TD)方法。为了进一步提高TD-FALCON在复杂多智能体领域的学习效率，本文提出了策略共享和邻近智能体机制两种策略。通过在交通控制问题域和羊群任务上的实验，我们证明了这些策略使TD-FALCON在复杂的多智能体域保持功能和适应性。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文本刊更多论文

Scaling Up Multi-agent Reinforcement Learning in Complex Domains

TD-FALCON (temporal difference-fusion architecture for learning, cognition, and navigation) is a class of self-organizing neural networks that incorporates temporal difference (TD) methods for real-time reinforcement learning. In this paper, we present two strategies, i.e. policy sharing and neighboring-agent mechanism, to further improve the learning efficiency of TD-FALCON in complex multi-agent domains. Through experiments on a traffic control problem domain and the herding task, we demonstrate that those strategies enable TD-FALCON to remain functional and adaptable in complex multi-agent domains.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

2008 IEEE/WIC/ACM International Conference on Web Intelligence and Intelligent Agent Technology

自引率

0.00%

发文量