基于改进异步优势参与者关键模型的不完全信息竞争策略

Proceedings of the 2020 4th International Conference on Deep Learning Technologies Pub Date : 2020-07-10 DOI:10.1145/3417188.3417189

Cong Zhao, Bing Xiao, Lin Zha

{"title":"基于改进异步优势参与者关键模型的不完全信息竞争策略","authors":"Cong Zhao, Bing Xiao, Lin Zha","doi":"10.1145/3417188.3417189","DOIUrl":null,"url":null,"abstract":"In recent years, game theory has been widely used in the field of deep learning, mainly including intelligent competition strategies of complete information games and incomplete information games. This paper focuses on incomplete information games, and proposes a low-dimensional semantic feature based on category coding and an incomplete information competition strategy based on the improved Asynchronous Advantage Actor-Critic (A3C) network model. First, the A3C network model in deep reinforcement learning is adopted in the competition strategy, and its network structure is improved according to the semantic features based on category coding. The improved A3C model is implemented in parallel by a series of \"workers\". The \"workers\" is a new deep learning model structure proposed in this paper. Secondly, this article combines supervised learning and Deep Reinforcement Learning (DRL) to propose a new competitive strategy. Through conducting a large number of real-time experiments with human players on online competitive websites, the comparison with the existing methods in terms of the ratio of winning and losing and the ranking rate, the experimental results indicate the superiority of the new competitive strategy.","PeriodicalId":373913,"journal":{"name":"Proceedings of the 2020 4th International Conference on Deep Learning Technologies","volume":"2015 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2020-07-10","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":"{\"title\":\"Incomplete Information Competition Strategy Based on Improved Asynchronous Advantage Actor Critical Model\",\"authors\":\"Cong Zhao, Bing Xiao, Lin Zha\",\"doi\":\"10.1145/3417188.3417189\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"In recent years, game theory has been widely used in the field of deep learning, mainly including intelligent competition strategies of complete information games and incomplete information games. This paper focuses on incomplete information games, and proposes a low-dimensional semantic feature based on category coding and an incomplete information competition strategy based on the improved Asynchronous Advantage Actor-Critic (A3C) network model. First, the A3C network model in deep reinforcement learning is adopted in the competition strategy, and its network structure is improved according to the semantic features based on category coding. The improved A3C model is implemented in parallel by a series of \\\"workers\\\". The \\\"workers\\\" is a new deep learning model structure proposed in this paper. Secondly, this article combines supervised learning and Deep Reinforcement Learning (DRL) to propose a new competitive strategy. Through conducting a large number of real-time experiments with human players on online competitive websites, the comparison with the existing methods in terms of the ratio of winning and losing and the ranking rate, the experimental results indicate the superiority of the new competitive strategy.\",\"PeriodicalId\":373913,\"journal\":{\"name\":\"Proceedings of the 2020 4th International Conference on Deep Learning Technologies\",\"volume\":\"2015 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2020-07-10\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"Proceedings of the 2020 4th International Conference on Deep Learning Technologies\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1145/3417188.3417189\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"Proceedings of the 2020 4th International Conference on Deep Learning Technologies","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1145/3417188.3417189","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 0

摘要

近年来，博弈论在深度学习领域得到了广泛的应用，主要包括完全信息博弈和不完全信息博弈的智能竞争策略。本文以不完全信息博弈为研究对象，提出了一种基于类别编码的低维语义特征和一种基于改进的异步优势参与者-批评者(A3C)网络模型的不完全信息竞争策略。首先，在竞争策略中采用深度强化学习中的A3C网络模型，并根据基于类别编码的语义特征对其网络结构进行改进。改进的A3C模型由一系列“工人”并行实现。“工人”是本文提出的一种新的深度学习模型结构。其次，本文将监督学习与深度强化学习(DRL)相结合，提出一种新的竞争策略。通过在在线竞技网站上对人类棋手进行大量的实时实验，并与现有方法在胜败比和排名率方面进行比较，实验结果表明了新竞技策略的优越性。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文本刊更多论文

Incomplete Information Competition Strategy Based on Improved Asynchronous Advantage Actor Critical Model

In recent years, game theory has been widely used in the field of deep learning, mainly including intelligent competition strategies of complete information games and incomplete information games. This paper focuses on incomplete information games, and proposes a low-dimensional semantic feature based on category coding and an incomplete information competition strategy based on the improved Asynchronous Advantage Actor-Critic (A3C) network model. First, the A3C network model in deep reinforcement learning is adopted in the competition strategy, and its network structure is improved according to the semantic features based on category coding. The improved A3C model is implemented in parallel by a series of "workers". The "workers" is a new deep learning model structure proposed in this paper. Secondly, this article combines supervised learning and Deep Reinforcement Learning (DRL) to propose a new competitive strategy. Through conducting a large number of real-time experiments with human players on online competitive websites, the comparison with the existing methods in terms of the ratio of winning and losing and the ranking rate, the experimental results indicate the superiority of the new competitive strategy.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

Proceedings of the 2020 4th International Conference on Deep Learning Technologies

自引率

0.00%

发文量