两相滤波在索引近似字符串匹配中的应用及其在唯一寡核苷酸搜索中的应用

Proceedings Eighth Symposium on String Processing and Information Retrieval Pub Date : 1900-01-01 DOI:10.1109/spire.2001.989742

H. Hyyro

{"title":"两相滤波在索引近似字符串匹配中的应用及其在唯一寡核苷酸搜索中的应用","authors":"H. Hyyro","doi":"10.1109/spire.2001.989742","DOIUrl":null,"url":null,"abstract":"We discuss using an indexing scheme to accelerate approximate search over a static text in the case of using unit cost edit distance as the measure of similarity between strings. First we generally consider the filtering criteria that can be used as a basis for the index, and then propose using filtering twice before the final checking phase. The last part consists of presenting an indexed approximate string matching application in bioinformatics, which is the search of unique oligonucleotides. We present practical comparisons and results for using different filtering schemes in this application. Our tests have involved a total of 15 different genomes, from which we present some results involving the largest two of these: The genome of Saccharomyces cerevisiae (baker's yeast) and a recent draft of the human genome, the latter being also the main target of the application.","PeriodicalId":107511,"journal":{"name":"Proceedings Eighth Symposium on String Processing and Information Retrieval","volume":"13 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"1900-01-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"3","resultStr":"{\"title\":\"On using two-phase filtering in indexed approximate string matching with application to searching unique oligonucleotides\",\"authors\":\"H. Hyyro\",\"doi\":\"10.1109/spire.2001.989742\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"We discuss using an indexing scheme to accelerate approximate search over a static text in the case of using unit cost edit distance as the measure of similarity between strings. First we generally consider the filtering criteria that can be used as a basis for the index, and then propose using filtering twice before the final checking phase. The last part consists of presenting an indexed approximate string matching application in bioinformatics, which is the search of unique oligonucleotides. We present practical comparisons and results for using different filtering schemes in this application. Our tests have involved a total of 15 different genomes, from which we present some results involving the largest two of these: The genome of Saccharomyces cerevisiae (baker's yeast) and a recent draft of the human genome, the latter being also the main target of the application.\",\"PeriodicalId\":107511,\"journal\":{\"name\":\"Proceedings Eighth Symposium on String Processing and Information Retrieval\",\"volume\":\"13 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"1900-01-01\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"3\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"Proceedings Eighth Symposium on String Processing and Information Retrieval\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1109/spire.2001.989742\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"Proceedings Eighth Symposium on String Processing and Information Retrieval","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/spire.2001.989742","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 3

摘要

我们讨论了在使用单位成本编辑距离作为字符串之间相似性度量的情况下，使用索引方案来加速对静态文本的近似搜索。首先我们一般考虑可以作为索引基础的过滤标准，然后建议在最后的检查阶段之前使用两次过滤。最后一部分介绍了一种基于索引的近似字符串匹配在生物信息学中的应用，即独特寡核苷酸的搜索。我们给出了在该应用中使用不同滤波方案的实际比较和结果。我们的测试总共涉及了15个不同的基因组，从中我们展示了其中最大的两个基因组的一些结果:酿酒酵母(面包酵母)的基因组和最近的人类基因组草图，后者也是该应用程序的主要目标。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文本刊更多论文

On using two-phase filtering in indexed approximate string matching with application to searching unique oligonucleotides

We discuss using an indexing scheme to accelerate approximate search over a static text in the case of using unit cost edit distance as the measure of similarity between strings. First we generally consider the filtering criteria that can be used as a basis for the index, and then propose using filtering twice before the final checking phase. The last part consists of presenting an indexed approximate string matching application in bioinformatics, which is the search of unique oligonucleotides. We present practical comparisons and results for using different filtering schemes in this application. Our tests have involved a total of 15 different genomes, from which we present some results involving the largest two of these: The genome of Saccharomyces cerevisiae (baker's yeast) and a recent draft of the human genome, the latter being also the main target of the application.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

Proceedings Eighth Symposium on String Processing and Information Retrieval

自引率

0.00%

发文量