Batched Nonparametric Contextual Bandits

IF 2.2 3区计算机科学 Q3 COMPUTER SCIENCE, INFORMATION SYSTEMS

IEEE Transactions on Information Theory Pub Date : 2025-03-26 DOI:10.1109/TIT.2025.3555071

Rong Jiang;Cong Ma

引用次数: 0

Abstract

We study nonparametric contextual bandits under batch constraints, where the expected reward for each action is modeled as a smooth function of covariates, and the policy updates are made at the end of each batch of observations. We establish a minimax regret lower bound for this setting and propose a novel batch learning algorithm that achieves the optimal regret (up to logarithmic factors). In essence, our procedure dynamically splits the covariate space into smaller bins, carefully aligning their widths with the batch size. Our theoretical results suggest that for mathematical framework of contextual bandit, a nearly constant number of policy updates can attain optimal regret in the fully online setting.

查看原文本刊更多论文

批处理非参数上下文强盗

我们研究了批约束下的非参数上下文盗匪，其中每个动作的预期奖励被建模为协变量的平滑函数，并且在每批观察结束时进行策略更新。我们为这种设置建立了最小最大遗憾下界，并提出了一种新的批量学习算法，该算法可以实现最优遗憾（高达对数因子）。本质上，我们的过程动态地将协变量空间分成更小的箱子，仔细地将它们的宽度与批大小对齐。我们的理论结果表明，对于上下文强盗的数学框架，在完全在线设置下，几乎恒定数量的政策更新可以达到最佳后悔。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

IEEE Transactions on Information Theory 工程技术-工程：电子与电气

CiteScore

5.70

自引率

20.00%

发文量

514

审稿时长

12 months

期刊介绍： The IEEE Transactions on Information Theory is a journal that publishes theoretical and experimental papers concerned with the transmission, processing, and utilization of information. The boundaries of acceptable subject matter are intentionally not sharply delimited. Rather, it is hoped that as the focus of research activity changes, a flexible policy will permit this Transactions to follow suit. Current appropriate topics are best reflected by recent Tables of Contents; they are summarized in the titles of editorial areas that appear on the inside front cover.