High Dimensional Regression on Serum Analytes

Italian journal of public health Pub Date : 2012-12-31 DOI:10.2427/8672

Yuanzhang Li, E. Schwarz, S. Bahn, R. Yolken, D. Niebuhr

{"title":"High Dimensional Regression on Serum Analytes","authors":"Yuanzhang Li, E. Schwarz, S. Bahn, R. Yolken, D. Niebuhr","doi":"10.2427/8672","DOIUrl":null,"url":null,"abstract":"Regression of high dimensional data is particularly difficult when the number of observations is limited. Principal Component Analysis, canonical correlation analysis and factor analysis are commonly used methods to reduce data dimensions, but usually cannot find the most significant linear combination. The goal is usually to find a particular partition of the space X consisting of all independent factors. In this paper, we propose an approach to high dimensional regression for applications where N>K or N<K, where N is the sample size, k is the dimension of space X. The approach starts by finding the most significant linear combination and one of the most insignificant directions to decompose the sample space into two subspaces and reduce the dimension. Further, we examine the contributions of individual variables to those most significant vectors by the coefficients of the combinations to reduce the total number of variables in the selected space without losing the power of the prediction. We use the proposed approach to determine the potential association of 51 serum analytes with schizophrenia using data derived from a case control study (n=208). Numerical results demonstrate that the proposed approach can significantly improve dimension reduction.","PeriodicalId":89162,"journal":{"name":"Italian journal of public health","volume":"9 1","pages":""},"PeriodicalIF":0.0000,"publicationDate":"2012-12-31","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"2","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Italian journal of public health","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.2427/8672","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 2

Abstract

Regression of high dimensional data is particularly difficult when the number of observations is limited. Principal Component Analysis, canonical correlation analysis and factor analysis are commonly used methods to reduce data dimensions, but usually cannot find the most significant linear combination. The goal is usually to find a particular partition of the space X consisting of all independent factors. In this paper, we propose an approach to high dimensional regression for applications where N>K or N

查看原文本刊更多论文

血清分析物的高维回归

当观测值有限时，高维数据的回归尤其困难。主成分分析、典型相关分析和因子分析是常用的数据降维方法，但往往找不到最显著的线性组合。目标通常是找到由所有独立因子组成的空间X的特定分区。本文针对N>K或N

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊