Shahan Derkarabetian, James Starrett, Marshal Hedin
{"title":"Using natural history to guide supervised machine learning for cryptic species delimitation with genetic data.","authors":"Shahan Derkarabetian, James Starrett, Marshal Hedin","doi":"10.1186/s12983-022-00453-0","DOIUrl":null,"url":null,"abstract":"<p><p>The diversity of biological and ecological characteristics of organisms, and the underlying genetic patterns and processes of speciation, makes the development of universally applicable genetic species delimitation methods challenging. Many approaches, like those incorporating the multispecies coalescent, sometimes delimit populations and overestimate species numbers. This issue is exacerbated in taxa with inherently high population structure due to low dispersal ability, and in cryptic species resulting from nonecological speciation. These taxa present a conundrum when delimiting species: analyses rely heavily, if not entirely, on genetic data which over split species, while other lines of evidence lump. We showcase this conundrum in the harvester Theromaster brunneus, a low dispersal taxon with a wide geographic distribution and high potential for cryptic species. Integrating morphology, mitochondrial, and sub-genomic (double-digest RADSeq and ultraconserved elements) data, we find high discordance across analyses and data types in the number of inferred species, with further evidence that multispecies coalescent approaches over split. We demonstrate the power of a supervised machine learning approach in effectively delimiting cryptic species by creating a \"custom\" training data set derived from a well-studied lineage with similar biological characteristics as Theromaster. This novel approach uses known taxa with particular biological characteristics to inform unknown taxa with similar characteristics, using modern computational tools ideally suited for species delimitation. The approach also considers the natural history of organisms to make more biologically informed species delimitation decisions, and in principle is broadly applicable for taxa across the tree of life.</p>","PeriodicalId":55142,"journal":{"name":"Frontiers in Zoology","volume":null,"pages":null},"PeriodicalIF":2.6000,"publicationDate":"2022-02-22","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8862334/pdf/","citationCount":"11","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Frontiers in Zoology","FirstCategoryId":"99","ListUrlMain":"https://doi.org/10.1186/s12983-022-00453-0","RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q1","JCRName":"ZOOLOGY","Score":null,"Total":0}
引用次数: 11
Abstract
The diversity of biological and ecological characteristics of organisms, and the underlying genetic patterns and processes of speciation, makes the development of universally applicable genetic species delimitation methods challenging. Many approaches, like those incorporating the multispecies coalescent, sometimes delimit populations and overestimate species numbers. This issue is exacerbated in taxa with inherently high population structure due to low dispersal ability, and in cryptic species resulting from nonecological speciation. These taxa present a conundrum when delimiting species: analyses rely heavily, if not entirely, on genetic data which over split species, while other lines of evidence lump. We showcase this conundrum in the harvester Theromaster brunneus, a low dispersal taxon with a wide geographic distribution and high potential for cryptic species. Integrating morphology, mitochondrial, and sub-genomic (double-digest RADSeq and ultraconserved elements) data, we find high discordance across analyses and data types in the number of inferred species, with further evidence that multispecies coalescent approaches over split. We demonstrate the power of a supervised machine learning approach in effectively delimiting cryptic species by creating a "custom" training data set derived from a well-studied lineage with similar biological characteristics as Theromaster. This novel approach uses known taxa with particular biological characteristics to inform unknown taxa with similar characteristics, using modern computational tools ideally suited for species delimitation. The approach also considers the natural history of organisms to make more biologically informed species delimitation decisions, and in principle is broadly applicable for taxa across the tree of life.
期刊介绍:
Frontiers in Zoology is an open access, peer-reviewed online journal publishing high quality research articles and reviews on all aspects of animal life.
As a biological discipline, zoology has one of the longest histories. Today it occasionally appears as though, due to the rapid expansion of life sciences, zoology has been replaced by more or less independent sub-disciplines amongst which exchange is often sparse. However, the recent advance of molecular methodology into "classical" fields of biology, and the development of theories that can explain phenomena on different levels of organisation, has led to a re-integration of zoological disciplines promoting a broader than usual approach to zoological questions. Zoology has re-emerged as an integrative discipline encompassing the most diverse aspects of animal life, from the level of the gene to the level of the ecosystem.
Frontiers in Zoology is the first open access journal focusing on zoology as a whole. It aims to represent and re-unite the various disciplines that look at animal life from different perspectives and at providing the basis for a comprehensive understanding of zoological phenomena on all levels of analysis. Frontiers in Zoology provides a unique opportunity to publish high quality research and reviews on zoological issues that will be internationally accessible to any reader at no cost.
The journal was initiated and is supported by the Deutsche Zoologische Gesellschaft, one of the largest national zoological societies with more than a century-long tradition in promoting high-level zoological research.