{"title":"Graph-based spatial segmentation of areal data","authors":"Vivien Goepp , Jan van de Kassteele","doi":"10.1016/j.csda.2023.107908","DOIUrl":null,"url":null,"abstract":"<div><p><span>Smoothing is often used to improve the readability and interpretability of noisy areal data. However, there are many instances where the underlying quantity is discontinuous. For such cases, specific methods are needed to estimate the piecewise constant spatial process. A well-known approach in this setting is to perform segmentation of the signal using the adjacency graph, such as the graph-based fused lasso. However, this method does not scale well to large graphs. A new method is introduced for piecewise constant spatial estimation that </span><em>(i)</em> is faster to compute on large graphs and <em>(ii)</em> yields sparser models than the fused lasso (for the same amount of regularization), resulting in estimates that are easier to interpret. The method is illustrated on simulated data and applied to real data on overweight prevalence in the Netherlands. Healthy and unhealthy zones are identified, which cannot be explained by demographic or socio-economic characteristics alone. The method is found capable of identifying such zones and can assist policymakers with their health improving strategies.</p></div>","PeriodicalId":55225,"journal":{"name":"Computational Statistics & Data Analysis","volume":null,"pages":null},"PeriodicalIF":1.5000,"publicationDate":"2023-12-12","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Computational Statistics & Data Analysis","FirstCategoryId":"100","ListUrlMain":"https://www.sciencedirect.com/science/article/pii/S0167947323002190","RegionNum":3,"RegionCategory":"数学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q3","JCRName":"COMPUTER SCIENCE, INTERDISCIPLINARY APPLICATIONS","Score":null,"Total":0}
引用次数: 0
Abstract
Smoothing is often used to improve the readability and interpretability of noisy areal data. However, there are many instances where the underlying quantity is discontinuous. For such cases, specific methods are needed to estimate the piecewise constant spatial process. A well-known approach in this setting is to perform segmentation of the signal using the adjacency graph, such as the graph-based fused lasso. However, this method does not scale well to large graphs. A new method is introduced for piecewise constant spatial estimation that (i) is faster to compute on large graphs and (ii) yields sparser models than the fused lasso (for the same amount of regularization), resulting in estimates that are easier to interpret. The method is illustrated on simulated data and applied to real data on overweight prevalence in the Netherlands. Healthy and unhealthy zones are identified, which cannot be explained by demographic or socio-economic characteristics alone. The method is found capable of identifying such zones and can assist policymakers with their health improving strategies.
期刊介绍:
Computational Statistics and Data Analysis (CSDA), an Official Publication of the network Computational and Methodological Statistics (CMStatistics) and of the International Association for Statistical Computing (IASC), is an international journal dedicated to the dissemination of methodological research and applications in the areas of computational statistics and data analysis. The journal consists of four refereed sections which are divided into the following subject areas:
I) Computational Statistics - Manuscripts dealing with: 1) the explicit impact of computers on statistical methodology (e.g., Bayesian computing, bioinformatics,computer graphics, computer intensive inferential methods, data exploration, data mining, expert systems, heuristics, knowledge based systems, machine learning, neural networks, numerical and optimization methods, parallel computing, statistical databases, statistical systems), and 2) the development, evaluation and validation of statistical software and algorithms. Software and algorithms can be submitted with manuscripts and will be stored together with the online article.
II) Statistical Methodology for Data Analysis - Manuscripts dealing with novel and original data analytical strategies and methodologies applied in biostatistics (design and analytic methods for clinical trials, epidemiological studies, statistical genetics, or genetic/environmental interactions), chemometrics, classification, data exploration, density estimation, design of experiments, environmetrics, education, image analysis, marketing, model free data exploration, pattern recognition, psychometrics, statistical physics, image processing, robust procedures.
[...]
III) Special Applications - [...]
IV) Annals of Statistical Data Science [...]