Subset Wavelet Trees

Jarno N. Alanko, E. Biagi, S. Puglisi, Jaakko Vuohtoniemi
{"title":"Subset Wavelet Trees","authors":"Jarno N. Alanko, E. Biagi, S. Puglisi, Jaakko Vuohtoniemi","doi":"10.4230/LIPIcs.SEA.2023.4","DOIUrl":null,"url":null,"abstract":"Given an alphabet Σ of σ = | Σ | symbols, a degenerate (or indeterminate) string X is a sequence X = X [0] , X [1] . . . , X [ n − 1] of n subsets of Σ. Since their introduction in the mid 70s, degenerate strings have been widely studied, with applications driven by their being a natural model for sequences in which there is a degree of uncertainty about the precise symbol at a given position, such as those arising in genomics and proteomics. In this paper we introduce a new data structural tool for degenerate strings, called the subset wavelet tree (SubsetWT). A SubsetWT supports two basic operations on degenerate strings: subset-rank( i, c ), which returns the number of subsets up to the i -th subset in the degenerate string that contain the symbol c ; and subset-select( i, c ), which returns the index in the degenerate string of the i -th subset that contains symbol c . These queries are analogs of rank and select queries that have been widely studied for ordinary strings. Via experiments in a real genomics application in which degenerate strings are fundamental, we show that subset wavelet trees are practical data structures, and in particular offer an attractive space-time tradeoff. Along the way we investigate data structures for supporting (normal) rank queries on base-4 and base-3 sequences, which may be of independent interest. Our C++ implementations of the data structures are available at https://github.com/jnalanko/SubsetWT .","PeriodicalId":9448,"journal":{"name":"Bulletin of the Society of Sea Water Science, Japan","volume":"23 1","pages":"4:1-4:14"},"PeriodicalIF":0.0000,"publicationDate":"2023-01-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"3","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Bulletin of the Society of Sea Water Science, Japan","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.4230/LIPIcs.SEA.2023.4","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
引用次数: 3

Abstract

Given an alphabet Σ of σ = | Σ | symbols, a degenerate (or indeterminate) string X is a sequence X = X [0] , X [1] . . . , X [ n − 1] of n subsets of Σ. Since their introduction in the mid 70s, degenerate strings have been widely studied, with applications driven by their being a natural model for sequences in which there is a degree of uncertainty about the precise symbol at a given position, such as those arising in genomics and proteomics. In this paper we introduce a new data structural tool for degenerate strings, called the subset wavelet tree (SubsetWT). A SubsetWT supports two basic operations on degenerate strings: subset-rank( i, c ), which returns the number of subsets up to the i -th subset in the degenerate string that contain the symbol c ; and subset-select( i, c ), which returns the index in the degenerate string of the i -th subset that contains symbol c . These queries are analogs of rank and select queries that have been widely studied for ordinary strings. Via experiments in a real genomics application in which degenerate strings are fundamental, we show that subset wavelet trees are practical data structures, and in particular offer an attractive space-time tradeoff. Along the way we investigate data structures for supporting (normal) rank queries on base-4 and base-3 sequences, which may be of independent interest. Our C++ implementations of the data structures are available at https://github.com/jnalanko/SubsetWT .
子集小波树
给定Σ = | Σ |符号的字母表Σ,简并(或不确定)字符串X是一个序列X = X [0], X[1]…, Σ的n个子集中的X [n−1]。自70年代中期引入以来,简并弦得到了广泛的研究,其应用是由于它们是序列的自然模型,其中在给定位置上的精确符号存在一定程度的不确定性,例如基因组学和蛋白质组学中出现的序列。本文介绍了一种新的数据结构工具,称为子集小波树(SubsetWT)。SubsetWT支持对退化字符串的两种基本操作:subset-rank(i, c),它返回退化字符串中包含符号c的第i个子集的子集数;还有subset-select(i, c),它返回包含符号c的第i个子集的退化字符串中的索引。这些查询类似于对普通字符串进行了广泛研究的rank和select查询。通过一个以简并字符串为基础的真实基因组学应用实验,我们证明了子集小波树是实用的数据结构,特别是提供了一个有吸引力的时空权衡。在此过程中,我们研究了支持以4为基数和以3为基数序列的(正常)排名查询的数据结构,这可能是独立的兴趣。我们的数据结构的c++实现可以在https://github.com/jnalanko/SubsetWT上获得。
本文章由计算机程序翻译,如有差异,请以英文原文为准。
求助全文
约1分钟内获得全文 求助全文
来源期刊
自引率
0.00%
发文量
0
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
确定
请完成安全验证×
copy
已复制链接
快去分享给好友吧!
我知道了
右上角分享
点击右上角分享
0
联系我们:info@booksci.cn Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。 Copyright © 2023 布克学术 All rights reserved.
京ICP备2023020795号-1
ghs 京公网安备 11010802042870号
Book学术文献互助
Book学术文献互助群
群 号:481959085
Book学术官方微信