Effect of look-ahead search depth in learning position evaluation functions for Othello using -greedy exploration

2007 IEEE Symposium on Computational Intelligence and Games Pub Date : 2007-04-01 DOI:10.1109/CIG.2007.368100

T. Runarsson, Egill Orn Jonsson

引用次数: 12

Abstract

This paper studies the effect of varying the depth of look-ahead for heuristic search in temporal difference (TD) learning and game playing. The acquisition position evaluation functions for the game of Othello is studied. The paper provides important insights into the strengths and weaknesses of using different search depths during learning when epsi-greedy exploration is applied. The main findings are that contrary to popular belief, for Othello, better playing strategies are found when TD learning is applied with lower look-ahead search depths

查看原文本刊更多论文

前瞻性搜索深度对奥赛罗位置评价函数学习的影响

本文研究了在时间差分学习和博弈中，改变前瞻深度对启发式搜索的影响。研究了奥赛罗博弈的获取位置评价函数。本文提供了重要的见解，在学习中使用不同的搜索深度时，epsi贪婪的探索应用的优缺点。主要发现是，与普遍看法相反，对于奥赛罗来说，当TD学习应用于较低的前瞻性搜索深度时，可以找到更好的游戏策略

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

2007 IEEE Symposium on Computational Intelligence and Games

自引率

0.00%

发文量