SnakeLines: integrated set of computational pipelines for sequencing reads.

IF 1.5 Q3 MATHEMATICAL & COMPUTATIONAL BIOLOGY
Journal of Integrative Bioinformatics Pub Date : 2023-08-21 eCollection Date: 2023-09-01 DOI:10.1515/jib-2022-0059
Jaroslav Budiš, Werner Krampl, Marcel Kucharík, Rastislav Hekel, Adrián Goga, Jozef Sitarčík, Michal Lichvár, Dávid Smol'ak, Miroslav Böhmer, Andrej Baláž, František Ďuriš, Juraj Gazdarica, Katarína Šoltys, Ján Turňa, Ján Radvánszky, Tomáš Szemes
{"title":"SnakeLines: integrated set of computational pipelines for sequencing reads.","authors":"Jaroslav Budiš, Werner Krampl, Marcel Kucharík, Rastislav Hekel, Adrián Goga, Jozef Sitarčík, Michal Lichvár, Dávid Smol'ak, Miroslav Böhmer, Andrej Baláž, František Ďuriš, Juraj Gazdarica, Katarína Šoltys, Ján Turňa, Ján Radvánszky, Tomáš Szemes","doi":"10.1515/jib-2022-0059","DOIUrl":null,"url":null,"abstract":"<p><p>With the rapid growth of massively parallel sequencing technologies, still more laboratories are utilising sequenced DNA fragments for genomic analyses. Interpretation of sequencing data is, however, strongly dependent on bioinformatics processing, which is often too demanding for clinicians and researchers without a computational background. Another problem represents the reproducibility of computational analyses across separated computational centres with inconsistent versions of installed libraries and bioinformatics tools. We propose an easily extensible set of computational pipelines, called SnakeLines, for processing sequencing reads; including mapping, assembly, variant calling, viral identification, transcriptomics, and metagenomics analysis. Individual steps of an analysis, along with methods and their parameters can be readily modified in a single configuration file. Provided pipelines are embedded in virtual environments that ensure isolation of required resources from the host operating system, rapid deployment, and reproducibility of analysis across different Unix-based platforms. SnakeLines is a powerful framework for the automation of bioinformatics analyses, with emphasis on a simple set-up, modifications, extensibility, and reproducibility. The framework is already routinely used in various research projects and their applications, especially in the Slovak national surveillance of SARS-CoV-2.</p>","PeriodicalId":53625,"journal":{"name":"Journal of Integrative Bioinformatics","volume":null,"pages":null},"PeriodicalIF":1.5000,"publicationDate":"2023-08-21","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10757078/pdf/","citationCount":"2","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Journal of Integrative Bioinformatics","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1515/jib-2022-0059","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"2023/9/1 0:00:00","PubModel":"eCollection","JCR":"Q3","JCRName":"MATHEMATICAL & COMPUTATIONAL BIOLOGY","Score":null,"Total":0}
引用次数: 2

Abstract

With the rapid growth of massively parallel sequencing technologies, still more laboratories are utilising sequenced DNA fragments for genomic analyses. Interpretation of sequencing data is, however, strongly dependent on bioinformatics processing, which is often too demanding for clinicians and researchers without a computational background. Another problem represents the reproducibility of computational analyses across separated computational centres with inconsistent versions of installed libraries and bioinformatics tools. We propose an easily extensible set of computational pipelines, called SnakeLines, for processing sequencing reads; including mapping, assembly, variant calling, viral identification, transcriptomics, and metagenomics analysis. Individual steps of an analysis, along with methods and their parameters can be readily modified in a single configuration file. Provided pipelines are embedded in virtual environments that ensure isolation of required resources from the host operating system, rapid deployment, and reproducibility of analysis across different Unix-based platforms. SnakeLines is a powerful framework for the automation of bioinformatics analyses, with emphasis on a simple set-up, modifications, extensibility, and reproducibility. The framework is already routinely used in various research projects and their applications, especially in the Slovak national surveillance of SARS-CoV-2.

SnakeLines:用于测序读数的集成计算管道集。
随着大规模并行测序技术的快速发展,越来越多的实验室开始利用测序 DNA 片段进行基因组分析。然而,测序数据的解读在很大程度上依赖于生物信息学处理,这对于没有计算背景的临床医生和研究人员来说往往要求过高。另一个问题是,不同的计算中心安装的文库和生物信息学工具版本不一致,计算分析的可重复性也存在问题。我们提出了一套易于扩展的计算管道,称为 "SnakeLines",用于处理测序读数,包括映射、组装、变异调用、病毒识别、转录组学和元基因组学分析。分析的各个步骤、方法及其参数可在单个配置文件中轻松修改。所提供的流水线被嵌入虚拟环境中,确保所需资源与主机操作系统隔离、快速部署以及在不同的 Unix 平台上进行分析的可重复性。SnakeLines 是一个功能强大的生物信息学自动化分析框架,强调简单的设置、修改、可扩展性和可重复性。该框架已在多个研究项目及其应用中得到常规使用,特别是在斯洛伐克的 SARS-CoV-2 国家监测中。
本文章由计算机程序翻译,如有差异,请以英文原文为准。
求助全文
约1分钟内获得全文 求助全文
来源期刊
Journal of Integrative Bioinformatics
Journal of Integrative Bioinformatics Medicine-Medicine (all)
CiteScore
3.10
自引率
5.30%
发文量
27
审稿时长
12 weeks
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
确定
请完成安全验证×
copy
已复制链接
快去分享给好友吧!
我知道了
右上角分享
点击右上角分享
0
联系我们:info@booksci.cn Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。 Copyright © 2023 布克学术 All rights reserved.
京ICP备2023020795号-1
ghs 京公网安备 11010802042870号
Book学术文献互助
Book学术文献互助群
群 号:481959085
Book学术官方微信