基于维基百科的语料库参考工具

Joint International Conference on Human-Centered Computer Environments Pub Date : 2012-03-08 DOI:10.1145/2160749.2160751

Jason Ginsburg

{"title":"基于维基百科的语料库参考工具","authors":"Jason Ginsburg","doi":"10.1145/2160749.2160751","DOIUrl":null,"url":null,"abstract":"This paper describes a dictionary-like reference tool that is designed to help users find information that is similar to what one would find in a dictionary when looking up a word, except that this information is extracted automatically from large corpora. For a particular vocabulary item, a user can view frequency information, part-of-speech distribution, word-forms, definitions, example paragraphs and collocations. All of this information is extracted automatically from corpora and most of this information is extracted from Wikipedia. Since Wikipedia is a massive corpus covering a diverse range of general topics, this information is probably very representative of how target words are used in general. This project has applications for English language teachers and learners, as well as for language researchers.","PeriodicalId":407345,"journal":{"name":"Joint International Conference on Human-Centered Computer Environments","volume":"31 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2012-03-08","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"1","resultStr":"{\"title\":\"A Wikipedia-based corpus reference tool\",\"authors\":\"Jason Ginsburg\",\"doi\":\"10.1145/2160749.2160751\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"This paper describes a dictionary-like reference tool that is designed to help users find information that is similar to what one would find in a dictionary when looking up a word, except that this information is extracted automatically from large corpora. For a particular vocabulary item, a user can view frequency information, part-of-speech distribution, word-forms, definitions, example paragraphs and collocations. All of this information is extracted automatically from corpora and most of this information is extracted from Wikipedia. Since Wikipedia is a massive corpus covering a diverse range of general topics, this information is probably very representative of how target words are used in general. This project has applications for English language teachers and learners, as well as for language researchers.\",\"PeriodicalId\":407345,\"journal\":{\"name\":\"Joint International Conference on Human-Centered Computer Environments\",\"volume\":\"31 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2012-03-08\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"1\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"Joint International Conference on Human-Centered Computer Environments\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1145/2160749.2160751\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"Joint International Conference on Human-Centered Computer Environments","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1145/2160749.2160751","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 1

摘要

本文描述了一个类似词典的参考工具，旨在帮助用户查找与查找单词时在字典中查找相似的信息，只是这些信息是自动从大型语料库中提取的。对于特定的词汇项，用户可以查看频率信息、词性分布、词形、定义、示例段落和搭配。所有这些信息都是自动从语料库中提取出来的，其中大部分信息是从维基百科中提取出来的。由于维基百科是一个庞大的语料库，涵盖了各种各样的一般主题，因此这些信息可能非常能代表目标单词的一般使用情况。该项目适用于英语教师和学习者，以及语言研究人员。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文本刊更多论文

A Wikipedia-based corpus reference tool

This paper describes a dictionary-like reference tool that is designed to help users find information that is similar to what one would find in a dictionary when looking up a word, except that this information is extracted automatically from large corpora. For a particular vocabulary item, a user can view frequency information, part-of-speech distribution, word-forms, definitions, example paragraphs and collocations. All of this information is extracted automatically from corpora and most of this information is extracted from Wikipedia. Since Wikipedia is a massive corpus covering a diverse range of general topics, this information is probably very representative of how target words are used in general. This project has applications for English language teachers and learners, as well as for language researchers.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

Joint International Conference on Human-Centered Computer Environments

自引率

0.00%

发文量