{"title":"Fundamental Q-learning Algorithm in Finding Optimal Policy","authors":"Canyu Sun","doi":"10.1109/ICSGEA.2017.84","DOIUrl":null,"url":null,"abstract":"Based on the off-Policy TD Control-Q learning, an agent is trained by reinforcement learning to find the optimal policy to reach the terminal state in the paper, which includes exploring the five factors affecting the learning efficiency and the results. To learn function Q to estimate the pros and cons of taking the current action, it must try every possible state and every alternative action and make a summery in the process of learning. Therefore, there are two main methods in the process of learning: exploration and utilization. Exploration is a method to try new action that is undiscovered and aim to discover better actions. Utilization is a method to adopt the optimal policy which taking actions according to the information discovered.","PeriodicalId":326442,"journal":{"name":"2017 International Conference on Smart Grid and Electrical Automation (ICSGEA)","volume":"86 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2017-05-27","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"16","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"2017 International Conference on Smart Grid and Electrical Automation (ICSGEA)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/ICSGEA.2017.84","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
引用次数: 16
Abstract
Based on the off-Policy TD Control-Q learning, an agent is trained by reinforcement learning to find the optimal policy to reach the terminal state in the paper, which includes exploring the five factors affecting the learning efficiency and the results. To learn function Q to estimate the pros and cons of taking the current action, it must try every possible state and every alternative action and make a summery in the process of learning. Therefore, there are two main methods in the process of learning: exploration and utilization. Exploration is a method to try new action that is undiscovered and aim to discover better actions. Utilization is a method to adopt the optimal policy which taking actions according to the information discovered.