{"title":"在创建用于预测蛋白质等电点值的学习集期间对二维电泳数据的过滤","authors":"Vladlen S. Skvortsov, A. Rybina","doi":"10.18097/bmcrm00162","DOIUrl":null,"url":null,"abstract":"A number of simple filters formulated from general considerations that take into account the peculiarities of the experiments as well as results obtained in 2D electrophoresis experiments are considered. These filters can be used for automated dataset formation and verification of learning of system for predicting protein isoelectric point values. These include: (i) filtering obvious errors introduced during initial database formation; (ii) selection of a known plausible range of values; (iii) selection of a single variant among various proteoforms; (iv) selection within a preset value of electrophoretic shift deviation, etc. Using a dataset combining data from 8 maps of Homo sapiens, Mus musculus, and Rattus norvegicus, the application of this set of filters improved the R2 value of predictions from 0.44 to 0.67.","PeriodicalId":286037,"journal":{"name":"Biomedical Chemistry: Research and Methods","volume":"5 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"1900-01-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":"{\"title\":\"The Filtration of 2D Electrophoresis Data During Creation of a Learning Set for Prediction of the Value of the Isoelectric Point of Proteins\",\"authors\":\"Vladlen S. Skvortsov, A. Rybina\",\"doi\":\"10.18097/bmcrm00162\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"A number of simple filters formulated from general considerations that take into account the peculiarities of the experiments as well as results obtained in 2D electrophoresis experiments are considered. These filters can be used for automated dataset formation and verification of learning of system for predicting protein isoelectric point values. These include: (i) filtering obvious errors introduced during initial database formation; (ii) selection of a known plausible range of values; (iii) selection of a single variant among various proteoforms; (iv) selection within a preset value of electrophoretic shift deviation, etc. Using a dataset combining data from 8 maps of Homo sapiens, Mus musculus, and Rattus norvegicus, the application of this set of filters improved the R2 value of predictions from 0.44 to 0.67.\",\"PeriodicalId\":286037,\"journal\":{\"name\":\"Biomedical Chemistry: Research and Methods\",\"volume\":\"5 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"1900-01-01\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"Biomedical Chemistry: Research and Methods\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.18097/bmcrm00162\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"Biomedical Chemistry: Research and Methods","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.18097/bmcrm00162","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
The Filtration of 2D Electrophoresis Data During Creation of a Learning Set for Prediction of the Value of the Isoelectric Point of Proteins
A number of simple filters formulated from general considerations that take into account the peculiarities of the experiments as well as results obtained in 2D electrophoresis experiments are considered. These filters can be used for automated dataset formation and verification of learning of system for predicting protein isoelectric point values. These include: (i) filtering obvious errors introduced during initial database formation; (ii) selection of a known plausible range of values; (iii) selection of a single variant among various proteoforms; (iv) selection within a preset value of electrophoretic shift deviation, etc. Using a dataset combining data from 8 maps of Homo sapiens, Mus musculus, and Rattus norvegicus, the application of this set of filters improved the R2 value of predictions from 0.44 to 0.67.