Computational Statistics最新文献_第7页

Overlapping coefficient in network-based semi-supervised clustering 基于网络的半监督聚类中的重叠系数

IF 1.3 4区数学

Computational Statistics Pub Date : 2024-02-19 DOI: 10.1007/s00180-024-01457-6

Claudio Conversano, Luca Frigau, Giulia Contu

{"title":"Overlapping coefficient in network-based semi-supervised clustering","authors":"Claudio Conversano, Luca Frigau, Giulia Contu","doi":"10.1007/s00180-024-01457-6","DOIUrl":"https://doi.org/10.1007/s00180-024-01457-6","url":null,"abstract":"Network-based Semi-Supervised Clustering (NeSSC) is a semi-supervised approach for clustering in the presence of an outcome variable. It uses a classification or regression model on resampled versions of the original data to produce a proximity matrix that indicates the magnitude of the similarity between pairs of observations measured with respect to the outcome. This matrix is transformed into a complex network on which a community detection algorithm is applied to search for underlying community structures which is a partition of the instances into highly homogeneous clusters to be evaluated in terms of the outcome. In this paper, we focus on the case the outcome variable to be used in NeSSC is numeric and propose an alternative selection criterion of the optimal partition based on a measure of overlapping between density curves as well as a penalization criterion which takes accounts for the number of clusters in a candidate partition. Next, we consider the performance of the proposed method for some artificial datasets and for 20 different real datasets and compare NeSSC with the other three popular methods of semi-supervised clustering with a numeric outcome. Results show that NeSSC with the overlapping criterion works particularly well when a reduced number of clusters are scattered localized.","PeriodicalId":55223,"journal":{"name":"Computational Statistics","volume":"18 1","pages":""},"PeriodicalIF":1.3,"publicationDate":"2024-02-19","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"139927826","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"数学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}

引用次数: 0

First exit and Dirichlet problem for the nonisotropic tempered $$alpha$$ -stable processes 非各向同性节制 $$alpha$$ 稳定过程的首次出口和德里赫特问题

IF 1.3 4区数学

Computational Statistics Pub Date : 2024-02-15 DOI: 10.1007/s00180-024-01462-9

Xing Liu, Weihua Deng

引用次数: 0

Some new invariant sum tests and MAD tests for the assessment of Benford’s law 用于评估本福德定律的一些新的不变量总和检验和 MAD 检验

IF 1.3 4区数学

Computational Statistics Pub Date : 2024-02-13 DOI: 10.1007/s00180-024-01463-8

Wolfgang Kössler, Hans-J. Lenz, Xing D. Wang

{"title":"Some new invariant sum tests and MAD tests for the assessment of Benford’s law","authors":"Wolfgang Kössler, Hans-J. Lenz, Xing D. Wang","doi":"10.1007/s00180-024-01463-8","DOIUrl":"https://doi.org/10.1007/s00180-024-01463-8","url":null,"abstract":"The Benford law is used world-wide for detecting non-conformance or data fraud of numerical data. It says that the significand of a data set from the universe is not uniformly, but logarithmically distributed. Especially, the first non-zero digit is One with an approximate probability of 0.3. There are several tests available for testing Benford, the best known are Pearson’s (chi ^2)-test, the Kolmogorov–Smirnov test and a modified version of the MAD-test. In the present paper we propose some tests, three of the four invariant sum tests are new and they are motivated by the sum invariance property of the Benford law. Two distance measures are investigated, Euclidean and Mahalanobis distance of the standardized sums to the orign. We use the significands corresponding to the first significant digit as well as the second significant digit, respectively. Moreover, we suggest inproved versions of the MAD-test and obtain critical values that are independent of the sample sizes. For illustration the tests are applied to specifically selected data sets where prior knowledge is available about being or not being Benford. Furthermore we discuss the role of truncation of distributions.","PeriodicalId":55223,"journal":{"name":"Computational Statistics","volume":"170 1","pages":""},"PeriodicalIF":1.3,"publicationDate":"2024-02-13","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"139766122","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"数学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}

引用次数: 0

Convergence of the CUSUM estimation for a mean shift in linear processes with random coefficients 具有随机系数的线性过程中均值移动的 CUSUM 估计的收敛性

IF 1.3 4区数学

Computational Statistics Pub Date : 2024-02-12 DOI: 10.1007/s00180-024-01465-6

Yi Wu, Wei Wang, Xuejun Wang

引用次数: 0

Analysis of estimating the Bayes rule for Gaussian mixture models with a specified missing-data mechanism 对具有指定缺失数据机制的高斯混合物模型贝叶斯规则的估计分析

IF 1.3 4区数学

Computational Statistics Pub Date : 2024-02-10 DOI: 10.1007/s00180-023-01447-0

引用次数: 0

Finite mixture of regression models for censored data based on the skew-t distribution 基于 skew-t 分布的删减数据有限混合回归模型

IF 1.3 4区数学

Computational Statistics Pub Date : 2024-02-10 DOI: 10.1007/s00180-024-01459-4

Jiwon Park, Dipak K. Dey, Víctor H. Lachos

{"title":"Finite mixture of regression models for censored data based on the skew-t distribution","authors":"Jiwon Park, Dipak K. Dey, Víctor H. Lachos","doi":"10.1007/s00180-024-01459-4","DOIUrl":"https://doi.org/10.1007/s00180-024-01459-4","url":null,"abstract":"Finite mixture models have been widely used to model and analyze data from heterogeneous populations. In practical scenarios, these types of data often confront upper and/or lower detection limits due to the constraints imposed by experimental apparatuses. Additional complexity arises when measures of each mixture component significantly deviate from the normal distribution, manifesting characteristics such as multimodality, asymmetry, and heavy-tailed behavior, simultaneously. This paper introduces a flexible model tailored for censored data to address these intricacies, leveraging the finite mixture of skew-t distributions. An Expectation Conditional Maximization Either (ECME) algorithm, is developed to efficiently derive parameter estimates by iteratively maximizing the observed data log-likelihood function. The algorithm has closed-form expressions at the E-step that rely on formulas for the mean and variance of truncated skew-t distributions. Moreover, a method based on general information principles is presented for approximating the asymptotic covariance matrix of the estimators. Results obtained from the analysis of both simulated and real datasets demonstrate the proposed method’s effectiveness.","PeriodicalId":55223,"journal":{"name":"Computational Statistics","volume":"38 1","pages":""},"PeriodicalIF":1.3,"publicationDate":"2024-02-10","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"139765979","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"数学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}

引用次数: 0

A simulation model to analyze the behavior of a faculty retirement plan: a case study in Mexico 分析教师退休计划行为的模拟模型：墨西哥案例研究

IF 1.3 4区数学

Computational Statistics Pub Date : 2024-02-09 DOI: 10.1007/s00180-024-01456-7

Marco Antonio Montufar-Benítez, Jaime Mora-Vargas, Carlos Arturo Soto-Campos, Gilberto Pérez-Lechuga, José Raúl Castro-Esparza

{"title":"A simulation model to analyze the behavior of a faculty retirement plan: a case study in Mexico","authors":"Marco Antonio Montufar-Benítez, Jaime Mora-Vargas, Carlos Arturo Soto-Campos, Gilberto Pérez-Lechuga, José Raúl Castro-Esparza","doi":"10.1007/s00180-024-01456-7","DOIUrl":"https://doi.org/10.1007/s00180-024-01456-7","url":null,"abstract":"The main goal in this study was to determine confidence intervals for average age, average seniority, and average money-savings, for faculty members in a university retirement system using a simulation model. The simulation—built-in Arena—considers age, seniority, and the probability of continuing in the institution as the main input random variables in the model. An annual interest rate of 7% and an average annual salary increase of 3% were considered. The scenario simulated consisted of the teacher and the university making contributions, the faculty 5% of his salary, and the university 5% of the teacher’s salary. Since the base salaries with which teachers join to university are variable, we considered a monthly salary of MXN 23 181.2, corresponding to full-time teachers with middle salaries. The results obtained by a simulation of 30 replicates showed that the confidence intervals for the average age at retirement were (55.0, 55.2) years, for the average seniority (22.1, 22.3) years, and for the average savings amount (329 795.2, 341 287.0) MXN. Moreover, the risk that a retiree of 62 years of age and more of 25 years of work, is alive after his savings runs out is approximately 98% and this happens at 64 years of age.","PeriodicalId":55223,"journal":{"name":"Computational Statistics","volume":"4 1","pages":""},"PeriodicalIF":1.3,"publicationDate":"2024-02-09","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"139765967","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"数学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}

引用次数: 0

Fitting concentric elliptical shapes under general model 一般模型下的同心椭圆形拟合

IF 1.3 4区数学

Computational Statistics Pub Date : 2024-02-09 DOI: 10.1007/s00180-024-01460-x

{"title":"Fitting concentric elliptical shapes under general model","authors":"","doi":"10.1007/s00180-024-01460-x","DOIUrl":"https://doi.org/10.1007/s00180-024-01460-x","url":null,"abstract":"<h3>Abstract</h3> Fitting concentric ellipses is a crucial yet challenging task in image processing, pattern recognition, and astronomy. To address this complexity, researchers have introduced simplified models by imposing geometric assumptions. These assumptions enable the linearization of the model through reparameterization, allowing for the extension of various fitting methods. However, these restrictive assumptions often fail to hold in real-world scenarios, limiting their practical applicability. In this work, we propose two novel estimators that relax these assumptions: the Least Squares method (LS) and the Gradient Algebraic Fit (GRAF). Since these methods are iterative, we provide numerical implementations and strategies for obtaining reliable initial guesses. Moreover, we employ perturbation theory to conduct a first-order analysis, deriving the leading terms of their Mean Squared Errors and their theoretical lower bounds. Our theoretical findings reveal that the GRAF is statistically efficient, while the LS method is not. We further validate our theoretical results and the performance of the proposed estimators through a series of numerical experiments on both real and synthetic data.","PeriodicalId":55223,"journal":{"name":"Computational Statistics","volume":"40 1","pages":""},"PeriodicalIF":1.3,"publicationDate":"2024-02-09","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"139765832","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"数学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}

引用次数: 0

Exploring local explanations of nonlinear models using animated linear projections 利用动画线性投影探索非线性模型的局部解释

IF 1.3 4区数学

Computational Statistics Pub Date : 2024-01-31 DOI: 10.1007/s00180-023-01453-2

Nicholas Spyrison, Dianne Cook, Przemyslaw Biecek

引用次数: 0

Semiparametric regression modelling of current status competing risks data: a Bayesian approach 现状竞争风险数据的半参数回归建模：一种贝叶斯方法

IF 1.3 4区数学

Computational Statistics Pub Date : 2024-01-31 DOI: 10.1007/s00180-024-01455-8

Pavithra Hariharan, P. G. Sankaran

{"title":"Semiparametric regression modelling of current status competing risks data: a Bayesian approach","authors":"Pavithra Hariharan, P. G. Sankaran","doi":"10.1007/s00180-024-01455-8","DOIUrl":"https://doi.org/10.1007/s00180-024-01455-8","url":null,"abstract":"The current status censoring takes place in survival analysis when the exact event times are not known, but each individual is monitored once for their survival status. The current status data often arise in medical research, from situations that involve multiple causes of failure. Examining current status competing risks data, commonly encountered in epidemiological studies and clinical trials, is more advantageous with Bayesian methods compared to conventional approaches. They excel in integrating prior knowledge with the observed data and delivering accurate results even with small samples. Inspired by these advantages, the present study is pioneering in introducing a Bayesian framework for both modelling and analysis of current status competing risks data together with covariates. By means of the proportional hazards model, estimation procedures for the regression parameters and cumulative incidence functions are established assuming appropriate prior distributions. The posterior computation is performed using an adaptive Metropolis–Hastings algorithm. Methods for comparing and validating models have been devised. An assessment of the finite sample characteristics of the estimators is conducted through simulation studies. Through the application of this Bayesian approach to prostate cancer clinical trial data, its practical efficacy is demonstrated.","PeriodicalId":55223,"journal":{"name":"Computational Statistics","volume":"37 1","pages":""},"PeriodicalIF":1.3,"publicationDate":"2024-01-31","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"139649048","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"数学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}

引用次数: 0