{"title":"基于学习者的词库词汇复杂度分析中的频度、离散度和抽象性:降维与识别","authors":"H. Zhang, Yuting Han, Xingzi Zhang, Liuran Cui","doi":"10.1080/09296174.2020.1782716","DOIUrl":null,"url":null,"abstract":"ABSTRACT The current study incorporated a number of lexical sophistication indices including frequency, dispersion and abstractness of words. A learner-based word bank (inclusive of a Chinese middle-school vocabulary list, a Chinese high-school vocabulary list and a Chinese college-English-test vocabulary list) was manually coded based on two existing corpora: Corpus of Contemporary American English (COCA) and British National Corpus (BNC). Indices of frequency, dispersion and abstractness of the word bank were analysed to shed light on the predetermined categorization of lexical sophistication among second language learners. Based on the principal component analysis, the results demonstrated that dispersion was a unique factor loaded on all entered eight variables while word frequency and abstractness were extracted by the same factor in the learner-based word bank. Moreover, a follow-up MANOVA analysis with post hoc comparisons showed that lexical sophistication indices in general produced pronounced differences among the three levels of word lists. More critically, dispersion was found to be the only significant indicator to differentiate the three levels of word lists. Discussion centred on the uniqueness of dispersion in lexical sophistication and the shared algorithm in frequency and abstractness.","PeriodicalId":45514,"journal":{"name":"Journal of Quantitative Linguistics","volume":"29 1","pages":"195 - 211"},"PeriodicalIF":0.7000,"publicationDate":"2020-07-09","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://sci-hub-pdf.com/10.1080/09296174.2020.1782716","citationCount":"1","resultStr":"{\"title\":\"Frequency, Dispersion and Abstractness in the Lexical Sophistication Analysis of A Learner-Based Word Bank: Dimensionality Reduction and Identification\",\"authors\":\"H. Zhang, Yuting Han, Xingzi Zhang, Liuran Cui\",\"doi\":\"10.1080/09296174.2020.1782716\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"ABSTRACT The current study incorporated a number of lexical sophistication indices including frequency, dispersion and abstractness of words. A learner-based word bank (inclusive of a Chinese middle-school vocabulary list, a Chinese high-school vocabulary list and a Chinese college-English-test vocabulary list) was manually coded based on two existing corpora: Corpus of Contemporary American English (COCA) and British National Corpus (BNC). Indices of frequency, dispersion and abstractness of the word bank were analysed to shed light on the predetermined categorization of lexical sophistication among second language learners. Based on the principal component analysis, the results demonstrated that dispersion was a unique factor loaded on all entered eight variables while word frequency and abstractness were extracted by the same factor in the learner-based word bank. Moreover, a follow-up MANOVA analysis with post hoc comparisons showed that lexical sophistication indices in general produced pronounced differences among the three levels of word lists. More critically, dispersion was found to be the only significant indicator to differentiate the three levels of word lists. Discussion centred on the uniqueness of dispersion in lexical sophistication and the shared algorithm in frequency and abstractness.\",\"PeriodicalId\":45514,\"journal\":{\"name\":\"Journal of Quantitative Linguistics\",\"volume\":\"29 1\",\"pages\":\"195 - 211\"},\"PeriodicalIF\":0.7000,\"publicationDate\":\"2020-07-09\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"https://sci-hub-pdf.com/10.1080/09296174.2020.1782716\",\"citationCount\":\"1\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"Journal of Quantitative Linguistics\",\"FirstCategoryId\":\"98\",\"ListUrlMain\":\"https://doi.org/10.1080/09296174.2020.1782716\",\"RegionNum\":2,\"RegionCategory\":\"文学\",\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"0\",\"JCRName\":\"LANGUAGE & LINGUISTICS\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"Journal of Quantitative Linguistics","FirstCategoryId":"98","ListUrlMain":"https://doi.org/10.1080/09296174.2020.1782716","RegionNum":2,"RegionCategory":"文学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"0","JCRName":"LANGUAGE & LINGUISTICS","Score":null,"Total":0}
Frequency, Dispersion and Abstractness in the Lexical Sophistication Analysis of A Learner-Based Word Bank: Dimensionality Reduction and Identification
ABSTRACT The current study incorporated a number of lexical sophistication indices including frequency, dispersion and abstractness of words. A learner-based word bank (inclusive of a Chinese middle-school vocabulary list, a Chinese high-school vocabulary list and a Chinese college-English-test vocabulary list) was manually coded based on two existing corpora: Corpus of Contemporary American English (COCA) and British National Corpus (BNC). Indices of frequency, dispersion and abstractness of the word bank were analysed to shed light on the predetermined categorization of lexical sophistication among second language learners. Based on the principal component analysis, the results demonstrated that dispersion was a unique factor loaded on all entered eight variables while word frequency and abstractness were extracted by the same factor in the learner-based word bank. Moreover, a follow-up MANOVA analysis with post hoc comparisons showed that lexical sophistication indices in general produced pronounced differences among the three levels of word lists. More critically, dispersion was found to be the only significant indicator to differentiate the three levels of word lists. Discussion centred on the uniqueness of dispersion in lexical sophistication and the shared algorithm in frequency and abstractness.
期刊介绍:
The Journal of Quantitative Linguistics is an international forum for the publication and discussion of research on the quantitative characteristics of language and text in an exact mathematical form. This approach, which is of growing interest, opens up important and exciting theoretical perspectives, as well as solutions for a wide range of practical problems such as machine learning or statistical parsing, by introducing into linguistics the methods and models of advanced scientific disciplines such as the natural sciences, economics, and psychology.