TE_Bench: a foundational benchmarking workflow for transposable element annotation pipelines.

IF 4.2 2区 生物学 Q1 GENETICS & HEREDITY
Hannah P Kania, Sierra A Seifert, Anne D Yoder
{"title":"TE_Bench: a foundational benchmarking workflow for transposable element annotation pipelines.","authors":"Hannah P Kania, Sierra A Seifert, Anne D Yoder","doi":"10.1186/s13100-026-00405-z","DOIUrl":null,"url":null,"abstract":"<p><strong>Background: </strong>Methods for computational discovery of transposable elements (TEs) in DNA sequences are under continual development. Although many TE annotation pipelines have been published, differences in their methods are often opaque and can lead to inconsistent results. As new methods emerge, there is a growing need for an informative and reproducible strategy to evaluate pipeline performance.</p><p><strong>Results: </strong>We developed TE_Bench, a user-friendly TE annotation benchmarking workflow that streamlines data generation and visualization for systematic comparison of annotation pipeline performance. TE_Bench can automate simulation of DNA sequences containing artificially evolved TEs from a database or accept user-provided real data to quantify how well test pipelines detect TEs relative to a reference annotation. To accommodate users with varying starting points, TE_Bench is housed as a Snakemake workflow with several options. The data it generates can be used to determine which TE annotation pipeline to use for a specific task, or to inform future improvements to pipelines by revealing shortcomings. We demonstrate the utility of TE_Bench in both contexts by benchmarking EDTA, RepeatModeler2, and Earl Grey using simulated and real DNA sequences. With simulated data, we assess the impact of variables including TE class and nested structure on annotation quality, and with real data, we consider whether read type and mapping method influence downstream annotation. In their default configurations, RepeatModeler2 and Earl Grey perform similarly, outperforming EDTA.</p><p><strong>Conclusions: </strong>TE_Bench is an extensible, open-source workflow that supports community-driven TE annotation benchmarking using simulated ground-truth or real genomes. TE_Bench can be accessed at https://gitub.com/hkania/TE_Bench.</p>","PeriodicalId":18854,"journal":{"name":"Mobile DNA","volume":" ","pages":""},"PeriodicalIF":4.2000,"publicationDate":"2026-07-16","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Mobile DNA","FirstCategoryId":"99","ListUrlMain":"https://doi.org/10.1186/s13100-026-00405-z","RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q1","JCRName":"GENETICS & HEREDITY","Score":null,"Total":0}
引用次数: 0

Abstract

Background: Methods for computational discovery of transposable elements (TEs) in DNA sequences are under continual development. Although many TE annotation pipelines have been published, differences in their methods are often opaque and can lead to inconsistent results. As new methods emerge, there is a growing need for an informative and reproducible strategy to evaluate pipeline performance.

Results: We developed TE_Bench, a user-friendly TE annotation benchmarking workflow that streamlines data generation and visualization for systematic comparison of annotation pipeline performance. TE_Bench can automate simulation of DNA sequences containing artificially evolved TEs from a database or accept user-provided real data to quantify how well test pipelines detect TEs relative to a reference annotation. To accommodate users with varying starting points, TE_Bench is housed as a Snakemake workflow with several options. The data it generates can be used to determine which TE annotation pipeline to use for a specific task, or to inform future improvements to pipelines by revealing shortcomings. We demonstrate the utility of TE_Bench in both contexts by benchmarking EDTA, RepeatModeler2, and Earl Grey using simulated and real DNA sequences. With simulated data, we assess the impact of variables including TE class and nested structure on annotation quality, and with real data, we consider whether read type and mapping method influence downstream annotation. In their default configurations, RepeatModeler2 and Earl Grey perform similarly, outperforming EDTA.

Conclusions: TE_Bench is an extensible, open-source workflow that supports community-driven TE annotation benchmarking using simulated ground-truth or real genomes. TE_Bench can be accessed at https://gitub.com/hkania/TE_Bench.

TE_Bench:可转置元素注释管道的基本基准测试工作流。
背景:DNA序列中转座因子(te)的计算发现方法正在不断发展。尽管已经发布了许多TE注释管道,但是它们的方法之间的差异通常是不透明的,并且可能导致不一致的结果。随着新方法的出现,人们越来越需要一种信息丰富、可重复的策略来评估管道性能。结果:我们开发了TE_Bench,一个用户友好的TE注释基准测试工作流,简化了数据生成和可视化,用于系统比较注释管道的性能。TE_Bench可以自动模拟数据库中包含人工进化的te的DNA序列,或者接受用户提供的真实数据,以量化测试管道相对于参考注释检测te的效果。为了适应不同起点的用户,TE_Bench被安置为具有多个选项的Snakemake工作流。它生成的数据可用于确定将哪个TE注释管道用于特定任务,或者通过揭示缺陷来通知管道的未来改进。我们通过使用模拟和真实的DNA序列对EDTA、RepeatModeler2和Earl Grey进行基准测试,展示了TE_Bench在这两种情况下的实用性。在模拟数据中,我们评估了TE类和嵌套结构等变量对标注质量的影响;在真实数据中,我们考虑了读取类型和映射方法对下游标注的影响。在其默认配置中,RepeatModeler2和Earl Grey的性能相似,优于EDTA。结论:TE_Bench是一个可扩展的开源工作流,支持社区驱动的TE注释基准测试,使用模拟的真实或真实基因组。TE_Bench可以通过https://gitub.com/hkania/TE_Bench访问。
本文章由计算机程序翻译,如有差异,请以英文原文为准。
求助全文
约1分钟内获得全文 求助全文
来源期刊
Mobile DNA
Mobile DNA GENETICS & HEREDITY-
CiteScore
8.20
自引率
6.10%
发文量
26
审稿时长
11 weeks
期刊介绍: Mobile DNA is an online, peer-reviewed, open access journal that publishes articles providing novel insights into DNA rearrangements in all organisms, ranging from transposition and other types of recombination mechanisms to patterns and processes of mobile element and host genome evolution. In addition, the journal will consider articles on the utility of mobile genetic elements in biotechnological methods and protocols.
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
确定
请完成安全验证×
copy
已复制链接
快去分享给好友吧!
我知道了
右上角分享
点击右上角分享
0
联系我们:info@booksci.cn Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。 Copyright © 2023 布克学术 All rights reserved.
京ICP备2023020795号-1
ghs 京公网安备 11010802042870号
Book学术文献互助
Book学术文献互助群
群 号:604180095
Book学术官方微信
小红书