Iyad Zimmo, Daniel Hörcher, Ramandeep Singh, Daniel J. Graham
{"title":"Benchmarking Travel Time and Demand Prediction Methods Using Large-scale Metro Smart Card Data","authors":"Iyad Zimmo, Daniel Hörcher, Ramandeep Singh, Daniel J. Graham","doi":"10.3311/pptr.22252","DOIUrl":null,"url":null,"abstract":"Urban mass transit systems generate large volumes of data via automated systems established for ticketing, signalling, and other operational processes. This study is motivated by the observation that despite the availability of sophisticated quantitative methods, most public transport operators are constrained in exploiting the information their datasets contain. This paper intends to address this gap in the context of real-time demand and travel time prediction with smart card data. We comparatively benchmark the predictive performance of four quantitative prediction methods: multivariate linear regression (MVLR) and semiparametric regression (SPR) widely used in the econometric literature, and random forest regression (RFR) and support vector machine regression (SVMR) from machine learning. We find that the SVMR and RFR methods are the most accurate in travel flow and travel time prediction, respectively. However, we also find that the SPR technique offers lower computation time at the expense of minor inefficiency in predictive power in comparison with the two machine learning methods.","PeriodicalId":39536,"journal":{"name":"Periodica Polytechnica Transportation Engineering","volume":"17 7 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2023-06-21","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Periodica Polytechnica Transportation Engineering","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.3311/pptr.22252","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q2","JCRName":"Engineering","Score":null,"Total":0}
引用次数: 0
Abstract
Urban mass transit systems generate large volumes of data via automated systems established for ticketing, signalling, and other operational processes. This study is motivated by the observation that despite the availability of sophisticated quantitative methods, most public transport operators are constrained in exploiting the information their datasets contain. This paper intends to address this gap in the context of real-time demand and travel time prediction with smart card data. We comparatively benchmark the predictive performance of four quantitative prediction methods: multivariate linear regression (MVLR) and semiparametric regression (SPR) widely used in the econometric literature, and random forest regression (RFR) and support vector machine regression (SVMR) from machine learning. We find that the SVMR and RFR methods are the most accurate in travel flow and travel time prediction, respectively. However, we also find that the SPR technique offers lower computation time at the expense of minor inefficiency in predictive power in comparison with the two machine learning methods.
期刊介绍:
Periodica Polytechnica is a publisher of the Budapest University of Technology and Economics. It publishes seven international journals (Architecture, Chemical Engineering, Civil Engineering, Electrical Engineering, Mechanical Engineering, Social and Management Sciences, Transportation Engineering). The journals have free electronic versions.