见树见木:随机森林的预测边缘

IF 1.7 2区文学 0 LANGUAGE & LINGUISTICS

Corpus Linguistics and Linguistic Theory Pub Date : 2023-03-28 DOI:10.1515/cllt-2022-0083

Lukas Sönning, Jason Grafmiller

{"title":"见树见木:随机森林的预测边缘","authors":"Lukas Sönning, Jason Grafmiller","doi":"10.1515/cllt-2022-0083","DOIUrl":null,"url":null,"abstract":"Abstract Classification trees and random forests offer a number of attractive features to corpus data analysts. However, the way in which these models are typically reported – a decision tree and/or set of variable importance scores – offers insufficient information if interest centers on the (form of) relationship between (multiple) predictors and the outcome. This paper develops predictive margins as an interpretative approach to ensemble techniques such as random forests. These are model summaries in the form of adjusted predictions, which provide a clearer picture of patterns in the data and allow us to query a model on potential nonlinear associations and interactions among predictor variables. The present paper outlines the general strategy for forming predictive margins and addresses methodological issues from an explicitly (corpus) linguistic perspective. For illustration, we use data on the English genitive alternation and provide an R package and code for their implementation.","PeriodicalId":45605,"journal":{"name":"Corpus Linguistics and Linguistic Theory","volume":"0 1","pages":""},"PeriodicalIF":1.7000,"publicationDate":"2023-03-28","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":"{\"title\":\"Seeing the wood for the trees: predictive margins for random forests\",\"authors\":\"Lukas Sönning, Jason Grafmiller\",\"doi\":\"10.1515/cllt-2022-0083\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"Abstract Classification trees and random forests offer a number of attractive features to corpus data analysts. However, the way in which these models are typically reported – a decision tree and/or set of variable importance scores – offers insufficient information if interest centers on the (form of) relationship between (multiple) predictors and the outcome. This paper develops predictive margins as an interpretative approach to ensemble techniques such as random forests. These are model summaries in the form of adjusted predictions, which provide a clearer picture of patterns in the data and allow us to query a model on potential nonlinear associations and interactions among predictor variables. The present paper outlines the general strategy for forming predictive margins and addresses methodological issues from an explicitly (corpus) linguistic perspective. For illustration, we use data on the English genitive alternation and provide an R package and code for their implementation.\",\"PeriodicalId\":45605,\"journal\":{\"name\":\"Corpus Linguistics and Linguistic Theory\",\"volume\":\"0 1\",\"pages\":\"\"},\"PeriodicalIF\":1.7000,\"publicationDate\":\"2023-03-28\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"Corpus Linguistics and Linguistic Theory\",\"FirstCategoryId\":\"98\",\"ListUrlMain\":\"https://doi.org/10.1515/cllt-2022-0083\",\"RegionNum\":2,\"RegionCategory\":\"文学\",\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"0\",\"JCRName\":\"LANGUAGE & LINGUISTICS\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"Corpus Linguistics and Linguistic Theory","FirstCategoryId":"98","ListUrlMain":"https://doi.org/10.1515/cllt-2022-0083","RegionNum":2,"RegionCategory":"文学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"0","JCRName":"LANGUAGE & LINGUISTICS","Score":null,"Total":0}

引用次数: 0

摘要

摘要分类树和随机森林为语料库数据分析提供了许多有吸引力的特征。然而，如果兴趣集中在(多个)预测因子和结果之间的关系(形式)上，这些模型的典型报告方式——决策树和/或可变重要性分数集——提供的信息不足。本文发展预测边际作为一种解释方法集成技术，如随机森林。这些是调整预测形式的模型摘要，它提供了数据模式的更清晰的图像，并允许我们查询预测变量之间潜在的非线性关联和相互作用的模型。本文概述了形成预测边缘的一般策略，并从明确(语料库)语言学的角度解决了方法论问题。为了说明这一点，我们使用了英语属格替换的数据，并提供了一个R包和实现它们的代码。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文本刊更多论文

Seeing the wood for the trees: predictive margins for random forests

Abstract Classification trees and random forests offer a number of attractive features to corpus data analysts. However, the way in which these models are typically reported – a decision tree and/or set of variable importance scores – offers insufficient information if interest centers on the (form of) relationship between (multiple) predictors and the outcome. This paper develops predictive margins as an interpretative approach to ensemble techniques such as random forests. These are model summaries in the form of adjusted predictions, which provide a clearer picture of patterns in the data and allow us to query a model on potential nonlinear associations and interactions among predictor variables. The present paper outlines the general strategy for forming predictive margins and addresses methodological issues from an explicitly (corpus) linguistic perspective. For illustration, we use data on the English genitive alternation and provide an R package and code for their implementation.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

Corpus Linguistics and Linguistic Theory Multiple-

CiteScore

4.20

自引率

12.50%

发文量

期刊介绍： Corpus Linguistics and Linguistic Theory (CLLT) is a peer-reviewed journal publishing high-quality original corpus-based research focusing on theoretically relevant issues in all core areas of linguistic research, or other recognized topic areas. It provides a forum for researchers from different theoretical backgrounds and different areas of interest that share a commitment to the systematic and exhaustive analysis of naturally occurring language. Contributions from all theoretical frameworks are welcome but they should be addressed at a general audience and thus be explicit about their assumptions and discovery procedures and provide sufficient theoretical background to be accessible to researchers from different frameworks. Topics Corpus Linguistics Quantitative Linguistics Phonology Morphology Semantics Syntax Pragmatics.