Assessing predictability of environmental time series with statistical and machine learning models

IF 1.7 3区环境科学与生态学 Q4 ENVIRONMENTAL SCIENCES

Environmetrics Pub Date : 2024-07-05 DOI:10.1002/env.2864

Matthew Bonas, Abhirup Datta, Christopher K. Wikle, Edward L. Boone, Faten S. Alamri, Bhava Vyasa Hari, Indulekha Kavila, Susan J. Simmons, Shannon M. Jarvis, Wesley S. Burr, Daniel E. Pagendam, Won Chang, Stefano Castruccio

{"title":"Assessing predictability of environmental time series with statistical and machine learning models","authors":"Matthew Bonas, Abhirup Datta, Christopher K. Wikle, Edward L. Boone, Faten S. Alamri, Bhava Vyasa Hari, Indulekha Kavila, Susan J. Simmons, Shannon M. Jarvis, Wesley S. Burr, Daniel E. Pagendam, Won Chang, Stefano Castruccio","doi":"10.1002/env.2864","DOIUrl":null,"url":null,"abstract":"<p>The ever increasing popularity of machine learning methods in virtually all areas of science, engineering and beyond is poised to put established statistical modeling approaches into question. Environmental statistics is no exception, as popular constructs such as neural networks and decision trees are now routinely used to provide forecasts of physical processes ranging from air pollution to meteorology. This presents both challenges and opportunities to the statistical community, which could contribute to the machine learning literature with a model-based approach with formal uncertainty quantification. Should, however, classical statistical methodologies be discarded altogether in environmental statistics, and should our contribution be focused on formalizing machine learning constructs? This work aims at providing some answers to this thought-provoking question with two time series case studies where selected models from both the statistical and machine learning literature are compared in terms of forecasting skills, uncertainty quantification and computational time. Relative merits of both class of approaches are discussed, and broad open questions are formulated as a baseline for a discussion on the topic.</p>","PeriodicalId":50512,"journal":{"name":"Environmetrics","volume":"36 1","pages":""},"PeriodicalIF":1.7000,"publicationDate":"2024-07-05","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Environmetrics","FirstCategoryId":"93","ListUrlMain":"https://onlinelibrary.wiley.com/doi/10.1002/env.2864","RegionNum":3,"RegionCategory":"环境科学与生态学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q4","JCRName":"ENVIRONMENTAL SCIENCES","Score":null,"Total":0}

引用次数: 0

Abstract

The ever increasing popularity of machine learning methods in virtually all areas of science, engineering and beyond is poised to put established statistical modeling approaches into question. Environmental statistics is no exception, as popular constructs such as neural networks and decision trees are now routinely used to provide forecasts of physical processes ranging from air pollution to meteorology. This presents both challenges and opportunities to the statistical community, which could contribute to the machine learning literature with a model-based approach with formal uncertainty quantification. Should, however, classical statistical methodologies be discarded altogether in environmental statistics, and should our contribution be focused on formalizing machine learning constructs? This work aims at providing some answers to this thought-provoking question with two time series case studies where selected models from both the statistical and machine learning literature are compared in terms of forecasting skills, uncertainty quantification and computational time. Relative merits of both class of approaches are discussed, and broad open questions are formulated as a baseline for a discussion on the topic.

查看原文本刊更多论文

利用统计和机器学习模型评估环境时间序列的可预测性

机器学习方法在科学、工程及其他几乎所有领域的应用日益普及，这将使既定的统计建模方法受到质疑。环境统计也不例外，因为神经网络和决策树等流行的结构现在已被常规用于提供从空气污染到气象学等物理过程的预测。这给统计界带来了挑战和机遇，统计界可以通过基于模型的方法和正式的不确定性量化，为机器学习文献做出贡献。然而，在环境统计中是否应该完全抛弃传统的统计方法，我们的贡献是否应该集中在机器学习构造的形式化上？这项工作旨在通过两个时间序列案例研究，从预测技能、不确定性量化和计算时间等方面对统计文献和机器学习文献中的选定模型进行比较，从而为这一发人深省的问题提供一些答案。讨论了这两类方法的相对优点，并提出了广泛的开放性问题，作为讨论该主题的基线。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

Environmetrics 环境科学-环境科学

CiteScore

2.90

自引率

17.60%

发文量

审稿时长

18-36 weeks

期刊介绍： Environmetrics, the official journal of The International Environmetrics Society (TIES), an Association of the International Statistical Institute, is devoted to the dissemination of high-quality quantitative research in the environmental sciences. The journal welcomes pertinent and innovative submissions from quantitative disciplines developing new statistical and mathematical techniques, methods, and theories that solve modern environmental problems. Articles must proffer substantive, new statistical or mathematical advances to answer important scientific questions in the environmental sciences, or must develop novel or enhanced statistical methodology with clear applications to environmental science. New methods should be illustrated with recent environmental data.