Yao-San Lin, Wan-Ni Cheng, C. Chen, Der-Chiang Li, Hung-Yu Chen
{"title":"生成合成样本以改进混合数值和分类属性的小样本学习","authors":"Yao-San Lin, Wan-Ni Cheng, C. Chen, Der-Chiang Li, Hung-Yu Chen","doi":"10.1109/IIAI-AAI.2019.00121","DOIUrl":null,"url":null,"abstract":"The small data learning issue has existed for over one hundred years (since 1908) when the Student's t-distribution was first developed. Few statistical tools can evaluate a population appropriately if the sample size is too small; small samples can be remedied through virtual sample generation (VSG) methods, which are widely used in industry and machine learning. However, most VSG methods were developed for data having only numerical attributes, very few studies have dealt with nominal attributes and cause domain estimation limitations. Therefore, this paper proposes a method that generates virtual samples based on the discrete degrees of nominal attributes, and then estimates the general population domains by fuzzy membership functions. A backpropagation neural network model and a support vector regression model are used to test the efficiency of the proposed method, while the Wilcoxon-sign test is used to test the difference with raw data sets. The result shows that the proposed method can reduce the mean absolute error and enhance classification accuracy by generating virtual samples that have nominal attributes.","PeriodicalId":136474,"journal":{"name":"2019 8th International Congress on Advanced Applied Informatics (IIAI-AAI)","volume":"15 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2019-07-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":"{\"title\":\"Generating Synthetic Samples to Improve Small Sample Learning with Mixed Numerical and Categorical Attributes\",\"authors\":\"Yao-San Lin, Wan-Ni Cheng, C. Chen, Der-Chiang Li, Hung-Yu Chen\",\"doi\":\"10.1109/IIAI-AAI.2019.00121\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"The small data learning issue has existed for over one hundred years (since 1908) when the Student's t-distribution was first developed. Few statistical tools can evaluate a population appropriately if the sample size is too small; small samples can be remedied through virtual sample generation (VSG) methods, which are widely used in industry and machine learning. However, most VSG methods were developed for data having only numerical attributes, very few studies have dealt with nominal attributes and cause domain estimation limitations. Therefore, this paper proposes a method that generates virtual samples based on the discrete degrees of nominal attributes, and then estimates the general population domains by fuzzy membership functions. A backpropagation neural network model and a support vector regression model are used to test the efficiency of the proposed method, while the Wilcoxon-sign test is used to test the difference with raw data sets. The result shows that the proposed method can reduce the mean absolute error and enhance classification accuracy by generating virtual samples that have nominal attributes.\",\"PeriodicalId\":136474,\"journal\":{\"name\":\"2019 8th International Congress on Advanced Applied Informatics (IIAI-AAI)\",\"volume\":\"15 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2019-07-01\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"2019 8th International Congress on Advanced Applied Informatics (IIAI-AAI)\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1109/IIAI-AAI.2019.00121\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"2019 8th International Congress on Advanced Applied Informatics (IIAI-AAI)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/IIAI-AAI.2019.00121","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
Generating Synthetic Samples to Improve Small Sample Learning with Mixed Numerical and Categorical Attributes
The small data learning issue has existed for over one hundred years (since 1908) when the Student's t-distribution was first developed. Few statistical tools can evaluate a population appropriately if the sample size is too small; small samples can be remedied through virtual sample generation (VSG) methods, which are widely used in industry and machine learning. However, most VSG methods were developed for data having only numerical attributes, very few studies have dealt with nominal attributes and cause domain estimation limitations. Therefore, this paper proposes a method that generates virtual samples based on the discrete degrees of nominal attributes, and then estimates the general population domains by fuzzy membership functions. A backpropagation neural network model and a support vector regression model are used to test the efficiency of the proposed method, while the Wilcoxon-sign test is used to test the difference with raw data sets. The result shows that the proposed method can reduce the mean absolute error and enhance classification accuracy by generating virtual samples that have nominal attributes.