{"title":"Neglected Heterogeneity, Simpson’s Paradox, and the Anatomy of Least Squares","authors":"R. Winkelmann","doi":"10.1515/jem-2023-0028","DOIUrl":null,"url":null,"abstract":"Abstract When a sample combines data from two or more groups, multivariate regression yields a matrix-weighted average of the group-specific coefficient vectors. However, it is possible that the weighted average of a specific coefficient falls outside the range of the group-specific coefficients, and it may even have a different sign compared to both group-level coefficients, a manifestation of Simpson’s paradox. The result of the combined regression is then prone to misinterpretation. The purpose of this paper is to raise awareness of this problem and to state conditions under which such non-convex weighting or sign reversal can arise, for a model with two regressors and two groups. Two illustrative examples, an investment equation estimated with panel data, and a cross-sectional earnings equation for men and women, highlight the relevance of these findings for applied work.","PeriodicalId":36727,"journal":{"name":"Journal of Econometric Methods","volume":"0 1","pages":""},"PeriodicalIF":0.0000,"publicationDate":"2023-08-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Journal of Econometric Methods","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1515/jem-2023-0028","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q3","JCRName":"Mathematics","Score":null,"Total":0}
引用次数: 0
Abstract
Abstract When a sample combines data from two or more groups, multivariate regression yields a matrix-weighted average of the group-specific coefficient vectors. However, it is possible that the weighted average of a specific coefficient falls outside the range of the group-specific coefficients, and it may even have a different sign compared to both group-level coefficients, a manifestation of Simpson’s paradox. The result of the combined regression is then prone to misinterpretation. The purpose of this paper is to raise awareness of this problem and to state conditions under which such non-convex weighting or sign reversal can arise, for a model with two regressors and two groups. Two illustrative examples, an investment equation estimated with panel data, and a cross-sectional earnings equation for men and women, highlight the relevance of these findings for applied work.