{"title":"Gaussian Sampling Approach to deal with Imbalanced Telemetry Datasets in Industrial Applications*","authors":"S. Galve, V. Puig, Xavier Vilajosana","doi":"10.1109/MED59994.2023.10185829","DOIUrl":null,"url":null,"abstract":"Practical implementation of data analytics in industrial environments has always been a problematic area because of data availability and quality. In this paper, a Gaussian sampling methodology is proposed to address the problem of imbalanced telemetry datasets that is one of the root causes that make modelling less reliable. By generating subsets that achieve homogeneous density distributions this problem is addressed. By comparing the impact of this method with the baseline case of random sampling, this paper aims to address this problem and propose a practical solution. A case study based on an industrial cooling device is used to assess and illustrate the proposed approach.","PeriodicalId":270226,"journal":{"name":"2023 31st Mediterranean Conference on Control and Automation (MED)","volume":"4 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2023-06-26","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"2023 31st Mediterranean Conference on Control and Automation (MED)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/MED59994.2023.10185829","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
引用次数: 0
Abstract
Practical implementation of data analytics in industrial environments has always been a problematic area because of data availability and quality. In this paper, a Gaussian sampling methodology is proposed to address the problem of imbalanced telemetry datasets that is one of the root causes that make modelling less reliable. By generating subsets that achieve homogeneous density distributions this problem is addressed. By comparing the impact of this method with the baseline case of random sampling, this paper aims to address this problem and propose a practical solution. A case study based on an industrial cooling device is used to assess and illustrate the proposed approach.