{"title":"Application Of L1- Regularization Approach In QSAR Problem. Linear Regression And Artificial Neural Networks","authors":"M. I. Berdnyk, A. B. Zakharov, V. Ivanov","doi":"10.17721/moca.2019.79-90","DOIUrl":null,"url":null,"abstract":"One of the primary tasks of analytical chemistry and QSAR/QSPR researches is building of prognostic regression equations based on descriptors sets. The one of the most important problems here is to decrease the number of descriptors in the initial descriptor set which is usually way too big. In current investigation the descriptor set is proposed to be reduced employing the least absolute shrinkage and selection operator (LASSO) approach. Decreased descriptor sets were used for calculations with application of the following QSAR/QSPR methods: ordinary least squares (OLS), the least absolute deviation (LAD) regressions and artificial neural networks (ANN). Contrary to aforementioned methods principal component regression (PCR) and partial least squares (PLS) approaches can produce solutions containing numerous descriptors. In this article we compared the viability of these two different descriptor handling ideologies in application to molecular chemical and physical properties prediction. From the obtained results it is possible to see that there are tasks for which PCR and PLS approaches can fail to produce accurate regression equations. At the same time, methods OLS and LAD that use small amount of descriptors can provide viable solutions for the same cases. It was shown that these small sets of descriptors selected with LASSO approach can be used in ANN to obtain models with even better internal validation characteristics.","PeriodicalId":18626,"journal":{"name":"Methods and Objects of Chemical Analysis","volume":null,"pages":null},"PeriodicalIF":0.7000,"publicationDate":"2019-01-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Methods and Objects of Chemical Analysis","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.17721/moca.2019.79-90","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q4","JCRName":"CHEMISTRY, ANALYTICAL","Score":null,"Total":0}
引用次数: 0
Abstract
One of the primary tasks of analytical chemistry and QSAR/QSPR researches is building of prognostic regression equations based on descriptors sets. The one of the most important problems here is to decrease the number of descriptors in the initial descriptor set which is usually way too big. In current investigation the descriptor set is proposed to be reduced employing the least absolute shrinkage and selection operator (LASSO) approach. Decreased descriptor sets were used for calculations with application of the following QSAR/QSPR methods: ordinary least squares (OLS), the least absolute deviation (LAD) regressions and artificial neural networks (ANN). Contrary to aforementioned methods principal component regression (PCR) and partial least squares (PLS) approaches can produce solutions containing numerous descriptors. In this article we compared the viability of these two different descriptor handling ideologies in application to molecular chemical and physical properties prediction. From the obtained results it is possible to see that there are tasks for which PCR and PLS approaches can fail to produce accurate regression equations. At the same time, methods OLS and LAD that use small amount of descriptors can provide viable solutions for the same cases. It was shown that these small sets of descriptors selected with LASSO approach can be used in ANN to obtain models with even better internal validation characteristics.
期刊介绍:
The journal "Methods and objects of chemical analysis" is peer-review journal and publishes original articles of theoretical and experimental analysis on topical issues and bio-analytical chemistry, chemical and pharmaceutical analysis, as well as chemical metrology. Submitted works shall cover the results of completed studies and shall make scientific contributions to the relevant area of expertise. The journal publishes review articles, research articles and articles related to latest developments of analytical instrumentations.