发布求助

文献互助智能选刊最新文献

Statistical properties of a class of randomized binary search algorithms

IF 0.8 4区计算机科学 Q4 COMPUTER SCIENCE, HARDWARE & ARCHITECTURE

Performance Evaluation Pub Date : 2025-03-05 DOI:10.1016/j.peva.2025.102478

Ye Xia

{"title":"Statistical properties of a class of randomized binary search algorithms","authors":"Ye Xia","doi":"10.1016/j.peva.2025.102478","DOIUrl":null,"url":null,"abstract":"<div><div>In this paper, we analyze the statistical properties of a randomized binary search algorithm and its variants. These algorithms have applications in caching and load balancing in distributed environments such as peer-to-peer networks, cloud storage, data centers, and content distribution networks. The basic discrete version of the problem is as follows. Suppose there are <math><mi>m</mi></math> servers, numbered 1, 2, …, <math><mi>m</mi></math>, out of which the first <math><mi>k</mi></math> servers are marked as special, where <math><mi>k</mi></math> is unknown. These <math><mi>k</mi></math> servers may contain a particular file or service that clients want. The objective is to select one of the marked servers uniformly at random. Considering the intended applications, we impose the constraint that there is no central controller to facilitate the selection process. We start with a basic algorithm: In each step, the client requesting the service chooses a number <math><mi>y</mi></math> uniformly at random from <math><mrow><mn>1</mn><mo>,</mo><mn>2</mn><mo>,</mo><mo>…</mo><mo>,</mo><mi>x</mi></mrow></math>, where <math><mi>x</mi></math> is the number chosen in the previous step, initially set to <math><mi>m</mi></math> in the first step. A query is then sent to server <math><mi>y</mi></math> asking whether <math><mi>y</mi></math> is marked. If the answer is yes, the algorithm returns <math><mi>y</mi></math>; otherwise, the process is repeated with <math><mrow><mi>x</mi><mo>←</mo><mi>y</mi></mrow></math>. In this paper, we primarily consider two batch versions of this algorithm in which multiple numbers are chosen in each step and multiple queries are made in parallel. We derive the mean and variance (exact and/or asymptotic) for the number of search steps in each version of the algorithm, and when possible, we give its distribution. Additionally, we analyze the access pattern of queries across the entire search space.</div></div>","PeriodicalId":19964,"journal":{"name":"Performance Evaluation","volume":"168 ","pages":"Article 102478"},"PeriodicalIF":0.8000,"publicationDate":"2025-03-05","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Performance Evaluation","FirstCategoryId":"94","ListUrlMain":"https://www.sciencedirect.com/science/article/pii/S0166531625000124","RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q4","JCRName":"COMPUTER SCIENCE, HARDWARE & ARCHITECTURE","Score":null,"Total":0}

引用次数: 0

Abstract

In this paper, we analyze the statistical properties of a randomized binary search algorithm and its variants. These algorithms have applications in caching and load balancing in distributed environments such as peer-to-peer networks, cloud storage, data centers, and content distribution networks. The basic discrete version of the problem is as follows. Suppose there are

m

servers, numbered 1, 2, …,

m

, out of which the first

k

servers are marked as special, where

k

is unknown. These

k

servers may contain a particular file or service that clients want. The objective is to select one of the marked servers uniformly at random. Considering the intended applications, we impose the constraint that there is no central controller to facilitate the selection process. We start with a basic algorithm: In each step, the client requesting the service chooses a number

y

uniformly at random from

1, 2, \dots, x

, where

x

is the number chosen in the previous step, initially set to

m

in the first step. A query is then sent to server

y

asking whether

y

is marked. If the answer is yes, the algorithm returns

y

; otherwise, the process is repeated with

x \leftarrow y

. In this paper, we primarily consider two batch versions of this algorithm in which multiple numbers are chosen in each step and multiple queries are made in parallel. We derive the mean and variance (exact and/or asymptotic) for the number of search steps in each version of the algorithm, and when possible, we give its distribution. Additionally, we analyze the access pattern of queries across the entire search space.

查看原文本刊更多论文

一类随机二叉搜索算法的统计性质

本文分析了一种随机化二分搜索算法及其变体的统计性质。这些算法在分布式环境（如点对点网络、云存储、数据中心和内容分发网络）中的缓存和负载平衡中有应用。这个问题的基本离散版本如下。假设有m个服务器，编号为1,2，…，m，其中前k个服务器被标记为特殊服务器，其中k是未知的。这些服务器可能包含客户端需要的特定文件或服务。。目标是均匀随机地选择一个标记的服务器。考虑到预期的应用程序，我们施加了没有中央控制器来促进选择过程的约束。我们从一个基本算法开始：在每一步中，请求服务的客户端从1、2、…、x中均匀随机选择一个数字y，其中x是在前一步中选择的数字，在第一步中初始设置为m。然后向服务器y发送查询，询问y是否被标记。如果答案是肯定的，算法返回y；否则，以x←y重复该过程。在本文中，我们主要考虑该算法的两个批处理版本，其中每一步选择多个数字，并行进行多个查询。我们推导出每个版本算法中搜索步骤数的均值和方差（精确和/或渐近），并在可能的情况下给出其分布。此外，我们还分析了整个搜索空间的查询访问模式。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

Performance Evaluation 工程技术-计算机：理论方法

CiteScore

3.10

自引率

0.00%

发文量

审稿时长

24 days

期刊介绍： Performance Evaluation functions as a leading journal in the area of modeling, measurement, and evaluation of performance aspects of computing and communication systems. As such, it aims to present a balanced and complete view of the entire Performance Evaluation profession. Hence, the journal is interested in papers that focus on one or more of the following dimensions: -Define new performance evaluation tools, including measurement and monitoring tools as well as modeling and analytic techniques -Provide new insights into the performance of computing and communication systems -Introduce new application areas where performance evaluation tools can play an important role and creative new uses for performance evaluation tools. More specifically, common application areas of interest include the performance of: -Resource allocation and control methods and algorithms (e.g. routing and flow control in networks, bandwidth allocation, processor scheduling, memory management) -System architecture, design and implementation -Cognitive radio -VANETs -Social networks and media -Energy efficient ICT -Energy harvesting -Data centers -Data centric networks -System reliability -System tuning and capacity planning -Wireless and sensor networks -Autonomic and self-organizing systems -Embedded systems -Network science