Tetsuya Nakatoh, Takahiko Suzuki, Tsukasa Kamimasu, S. Hirokawa
{"title":"统计数据非自然部分的检测","authors":"Tetsuya Nakatoh, Takahiko Suzuki, Tsukasa Kamimasu, S. Hirokawa","doi":"10.52731/iee.v6.i2.569","DOIUrl":null,"url":null,"abstract":"Ensuring the authenticity of statistical data is important because such data are used for various decision-making tasks. However, in practical applications, several types of data alterations have been reported. Therefore, it is necessary to validate the accuracy of statistical data. Benford’s law is a well-known method for detecting unnatural numerical data. According to Benford’s law, the occurrence probability of the first significant digits follows a particular distribution. However, the unnatural parts of data cannot be accurately identi-fied. In this study, we attempted to identify the unnatural parts of statistical data available in tabular format. A subset of the target data was specified using the row and column names that define each cell in the table or the words displayed in the table title. By measuring the divergence of the subsets, we identified the unnatural subsets. In this paper, we present the results of the identification of unnatural subsets using the agricultural data acquired from the China Statistical Yearbook.","PeriodicalId":416504,"journal":{"name":"Information Engineering Express","volume":"3 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2020-12-30","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":"{\"title\":\"Detection of Unnatural Parts of Statistical Data\",\"authors\":\"Tetsuya Nakatoh, Takahiko Suzuki, Tsukasa Kamimasu, S. Hirokawa\",\"doi\":\"10.52731/iee.v6.i2.569\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"Ensuring the authenticity of statistical data is important because such data are used for various decision-making tasks. However, in practical applications, several types of data alterations have been reported. Therefore, it is necessary to validate the accuracy of statistical data. Benford’s law is a well-known method for detecting unnatural numerical data. According to Benford’s law, the occurrence probability of the first significant digits follows a particular distribution. However, the unnatural parts of data cannot be accurately identi-fied. In this study, we attempted to identify the unnatural parts of statistical data available in tabular format. A subset of the target data was specified using the row and column names that define each cell in the table or the words displayed in the table title. By measuring the divergence of the subsets, we identified the unnatural subsets. In this paper, we present the results of the identification of unnatural subsets using the agricultural data acquired from the China Statistical Yearbook.\",\"PeriodicalId\":416504,\"journal\":{\"name\":\"Information Engineering Express\",\"volume\":\"3 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2020-12-30\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"Information Engineering Express\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.52731/iee.v6.i2.569\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"Information Engineering Express","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.52731/iee.v6.i2.569","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
Ensuring the authenticity of statistical data is important because such data are used for various decision-making tasks. However, in practical applications, several types of data alterations have been reported. Therefore, it is necessary to validate the accuracy of statistical data. Benford’s law is a well-known method for detecting unnatural numerical data. According to Benford’s law, the occurrence probability of the first significant digits follows a particular distribution. However, the unnatural parts of data cannot be accurately identi-fied. In this study, we attempted to identify the unnatural parts of statistical data available in tabular format. A subset of the target data was specified using the row and column names that define each cell in the table or the words displayed in the table title. By measuring the divergence of the subsets, we identified the unnatural subsets. In this paper, we present the results of the identification of unnatural subsets using the agricultural data acquired from the China Statistical Yearbook.