评估加利福尼亚州交通事故预测中可解释机器学习模型的性能

2020 39th International Conference of the Chilean Computer Science Society (SCCC) Pub Date : 2020-11-16 DOI:10.1109/SCCC51225.2020.9281196

Camilo Parra, C. Ponce, Rodrigo F. Salas

{"title":"评估加利福尼亚州交通事故预测中可解释机器学习模型的性能","authors":"Camilo Parra, C. Ponce, Rodrigo F. Salas","doi":"10.1109/SCCC51225.2020.9281196","DOIUrl":null,"url":null,"abstract":"Reducing and preventing road traffic accidents is a major public health problem and a priority for many nations. In this paper, we seek to explore the performance of explainable machine learning models applied to the prediction of road traffic crashes using a dataset containing nearly three million records of this type of events and the conditions under which they occurred. To achieve this, the dataset US Accidents -A Countrywide Traffic Accident Dataset is used. First we will clean, standardize and reduce the data, then we will transform the time and location values using a geohashing library developed by Uber, later, we will increase our dataset to obtain events classified as ‘not an accident’ using web scraping techniques in the data sources of the original authors of the dataset. Then, we will evaluate the performance of different implementations of Random Forest and decision trees, we obtained a performance superior to 70% for the F1 score of these models. Finally, we conclude that weather conditions are strongly related to the car accident.","PeriodicalId":117157,"journal":{"name":"2020 39th International Conference of the Chilean Computer Science Society (SCCC)","volume":"68 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2020-11-16","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"6","resultStr":"{\"title\":\"Evaluating the Performance of Explainable Machine Learning Models in Traffic Accidents Prediction in California\",\"authors\":\"Camilo Parra, C. Ponce, Rodrigo F. Salas\",\"doi\":\"10.1109/SCCC51225.2020.9281196\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"Reducing and preventing road traffic accidents is a major public health problem and a priority for many nations. In this paper, we seek to explore the performance of explainable machine learning models applied to the prediction of road traffic crashes using a dataset containing nearly three million records of this type of events and the conditions under which they occurred. To achieve this, the dataset US Accidents -A Countrywide Traffic Accident Dataset is used. First we will clean, standardize and reduce the data, then we will transform the time and location values using a geohashing library developed by Uber, later, we will increase our dataset to obtain events classified as ‘not an accident’ using web scraping techniques in the data sources of the original authors of the dataset. Then, we will evaluate the performance of different implementations of Random Forest and decision trees, we obtained a performance superior to 70% for the F1 score of these models. Finally, we conclude that weather conditions are strongly related to the car accident.\",\"PeriodicalId\":117157,\"journal\":{\"name\":\"2020 39th International Conference of the Chilean Computer Science Society (SCCC)\",\"volume\":\"68 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2020-11-16\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"6\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"2020 39th International Conference of the Chilean Computer Science Society (SCCC)\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1109/SCCC51225.2020.9281196\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"2020 39th International Conference of the Chilean Computer Science Society (SCCC)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/SCCC51225.2020.9281196","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 6

摘要

减少和预防道路交通事故是一个重大的公共卫生问题，也是许多国家的优先事项。在本文中，我们试图探索应用于道路交通碰撞预测的可解释机器学习模型的性能，使用包含近300万条此类事件记录及其发生条件的数据集。为了实现这一点，使用了美国事故数据集-全国交通事故数据集。首先，我们将清理、标准化和减少数据，然后我们将使用Uber开发的地理哈希库转换时间和位置值，之后，我们将增加我们的数据集，使用数据集原作者的数据源中的web抓取技术来获得分类为“非事故”的事件。然后，我们将评估随机森林和决策树的不同实现的性能，我们获得了这些模型的F1分数优于70%的性能。最后，我们得出结论，天气条件与车祸密切相关。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文本刊更多论文

Evaluating the Performance of Explainable Machine Learning Models in Traffic Accidents Prediction in California

Reducing and preventing road traffic accidents is a major public health problem and a priority for many nations. In this paper, we seek to explore the performance of explainable machine learning models applied to the prediction of road traffic crashes using a dataset containing nearly three million records of this type of events and the conditions under which they occurred. To achieve this, the dataset US Accidents -A Countrywide Traffic Accident Dataset is used. First we will clean, standardize and reduce the data, then we will transform the time and location values using a geohashing library developed by Uber, later, we will increase our dataset to obtain events classified as ‘not an accident’ using web scraping techniques in the data sources of the original authors of the dataset. Then, we will evaluate the performance of different implementations of Random Forest and decision trees, we obtained a performance superior to 70% for the F1 score of these models. Finally, we conclude that weather conditions are strongly related to the car accident.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

2020 39th International Conference of the Chilean Computer Science Society (SCCC)

自引率

0.00%

发文量