Coding Structure for the ORF1ab, S, M and N Coronavirus Genes

Q3 Mathematics
M. Chaley, Zh.S. Tyulko, V. Kutyrkin
{"title":"Coding Structure for the ORF1ab, S, M and N Coronavirus Genes","authors":"M. Chaley, Zh.S. Tyulko, V. Kutyrkin","doi":"10.17537/2020.15.441","DOIUrl":null,"url":null,"abstract":"\nSpectral-statistical approach was applied to comparative analysis of coronavirus genomes from the four genus Alphacoronavirus, Betacoronavirus (including new SARS-CoV-2 virus), Gammacoronavirus and Deltacoronavirus. This analysis was done from the point of view of 3-regularity and latent triplet profile periodicity existence in the coding sequences of four structural genes: ORF1ab encoding transcriptase; S-gene of glycoprotein forming spikes; M-gene of membrane protein; N-gene of nucleoprotein. A whole number of the genomes analyzed was equal to 3410. Gene numbers in each of the four groups in the study respectively were the same. In the result, practically, in the CDSs of all analyzed genes of ORF1ab, S and N the latent profile triplet periodicity was revealed and high value of 3-regularity index, being a quality estimate of coding triplet structure conservation, was determined. On the contrary, for coding structure of M-genes a tendency was revealed to diffuse up to homogeneity for 60 % of the genes in the genomes of alphacoronaviruses analyzed and for 67 % of the genes of the gammacoronaviruses. Tendency of the such structure diffusion, being accompanied by decrease of 3-regularity index average value in comparison with other genes, while the triplet profile periodicity remains saved, was also noted for M-genes of SARS-CoV-2 viruses. Probably, this tendency reflects a significance of M-genes variability in coronavirus adaptation to the novel hosts of genus. Analysis of 3-profile periodicity matrices of the four groups of SARS-CoV-2 genes considered in the work, for the viruses isolated in Europe, Asia and USA, did not revealed their significant difference, that is allowing to propose a single source of this virus propagation. \n","PeriodicalId":53525,"journal":{"name":"Mathematical Biology and Bioinformatics","volume":"1 1","pages":""},"PeriodicalIF":0.0000,"publicationDate":"2020-12-25","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"1","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Mathematical Biology and Bioinformatics","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.17537/2020.15.441","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q3","JCRName":"Mathematics","Score":null,"Total":0}
引用次数: 1

Abstract

Spectral-statistical approach was applied to comparative analysis of coronavirus genomes from the four genus Alphacoronavirus, Betacoronavirus (including new SARS-CoV-2 virus), Gammacoronavirus and Deltacoronavirus. This analysis was done from the point of view of 3-regularity and latent triplet profile periodicity existence in the coding sequences of four structural genes: ORF1ab encoding transcriptase; S-gene of glycoprotein forming spikes; M-gene of membrane protein; N-gene of nucleoprotein. A whole number of the genomes analyzed was equal to 3410. Gene numbers in each of the four groups in the study respectively were the same. In the result, practically, in the CDSs of all analyzed genes of ORF1ab, S and N the latent profile triplet periodicity was revealed and high value of 3-regularity index, being a quality estimate of coding triplet structure conservation, was determined. On the contrary, for coding structure of M-genes a tendency was revealed to diffuse up to homogeneity for 60 % of the genes in the genomes of alphacoronaviruses analyzed and for 67 % of the genes of the gammacoronaviruses. Tendency of the such structure diffusion, being accompanied by decrease of 3-regularity index average value in comparison with other genes, while the triplet profile periodicity remains saved, was also noted for M-genes of SARS-CoV-2 viruses. Probably, this tendency reflects a significance of M-genes variability in coronavirus adaptation to the novel hosts of genus. Analysis of 3-profile periodicity matrices of the four groups of SARS-CoV-2 genes considered in the work, for the viruses isolated in Europe, Asia and USA, did not revealed their significant difference, that is allowing to propose a single source of this virus propagation.
ORF1ab、S、M和N冠状病毒基因的编码结构
采用光谱统计方法对甲型冠状病毒、倍冠状病毒(包括新型SARS-CoV-2病毒)、伽玛冠状病毒和德尔塔冠状病毒4个属的冠状病毒基因组进行比较分析。从的角度分析做了3-regularity和潜在的三联体周期性存在四个结构基因的编码序列:ORF1ab编码转录酶;糖蛋白形成穗的s基因;膜蛋白m基因;核蛋白的n基因。被分析的基因组总数为3410个。研究中四组的基因数量都是相同的。结果表明,在ORF1ab、S和N基因的cds分析中,发现了潜在的三联体结构周期性,并确定了3-正则指数的高值,作为编码三联体结构保守性的质量评价。相反,对于m基因的编码结构,在所分析的甲型冠状病毒基因组中,60%的基因和67%的伽玛冠状病毒基因组中,有扩散到同质性的趋势。SARS-CoV-2病毒的m基因也有这种结构扩散的趋势,与其他基因相比,3-正则指数平均值降低,而三联体谱的周期性保持不变。这种趋势可能反映了m基因变异在冠状病毒适应新宿主中的重要意义。在欧洲、亚洲和美国分离的病毒中,对工作中考虑的四组SARS-CoV-2基因的3-profile周期性矩阵进行分析,未发现它们的显著差异,这允许提出该病毒传播的单一来源。
本文章由计算机程序翻译,如有差异,请以英文原文为准。
求助全文
约1分钟内获得全文 求助全文
来源期刊
Mathematical Biology and Bioinformatics
Mathematical Biology and Bioinformatics Mathematics-Applied Mathematics
CiteScore
1.10
自引率
0.00%
发文量
13
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
确定
请完成安全验证×
copy
已复制链接
快去分享给好友吧!
我知道了
右上角分享
点击右上角分享
0
联系我们:info@booksci.cn Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。 Copyright © 2023 布克学术 All rights reserved.
京ICP备2023020795号-1
ghs 京公网安备 11010802042870号
Book学术文献互助
Book学术文献互助群
群 号:481959085
Book学术官方微信