藏文新有效词自动提取方法

2012 Fifth International Conference on Intelligent Networks and Intelligent Systems Pub Date : 2012-11-01 DOI:10.1109/ICINIS.2012.61

Yuan Sun, Xiaodong Yan, Xiaobing Zhao, Guosheng Yang

{"title":"藏文新有效词自动提取方法","authors":"Yuan Sun, Xiaodong Yan, Xiaobing Zhao, Guosheng Yang","doi":"10.1109/ICINIS.2012.61","DOIUrl":null,"url":null,"abstract":"This paper proposes a model to automatically extract Tibetan new valid words. Through building the dynamic Tibetan corpus from 2009 to 2012, which covers more than 18 Tibetan network media of Tibet, Qinghai, Sichuan, Gansu and Yunnan, we research on the key techniques of Tibetan new valid word extraction: (1) using statistical method to establish Tibetan new word knowledge base, (2) using information entropy to filter Tibetan new valid words, (3) using vector space module similarity calculation to extract Tibetan new valid word.","PeriodicalId":302503,"journal":{"name":"2012 Fifth International Conference on Intelligent Networks and Intelligent Systems","volume":"4 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2012-11-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"1","resultStr":"{\"title\":\"Automatic Extraction Method of Tibetan New Valid Words\",\"authors\":\"Yuan Sun, Xiaodong Yan, Xiaobing Zhao, Guosheng Yang\",\"doi\":\"10.1109/ICINIS.2012.61\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"This paper proposes a model to automatically extract Tibetan new valid words. Through building the dynamic Tibetan corpus from 2009 to 2012, which covers more than 18 Tibetan network media of Tibet, Qinghai, Sichuan, Gansu and Yunnan, we research on the key techniques of Tibetan new valid word extraction: (1) using statistical method to establish Tibetan new word knowledge base, (2) using information entropy to filter Tibetan new valid words, (3) using vector space module similarity calculation to extract Tibetan new valid word.\",\"PeriodicalId\":302503,\"journal\":{\"name\":\"2012 Fifth International Conference on Intelligent Networks and Intelligent Systems\",\"volume\":\"4 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2012-11-01\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"1\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"2012 Fifth International Conference on Intelligent Networks and Intelligent Systems\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1109/ICINIS.2012.61\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"2012 Fifth International Conference on Intelligent Networks and Intelligent Systems","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/ICINIS.2012.61","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 1

摘要

本文提出了一种自动提取藏文新有效词的模型。通过构建2009 - 2012年西藏、青海、四川、甘肃、云南等18种藏语网络媒体的动态藏语语料库，研究了藏语新有效词提取的关键技术:(1)采用统计方法建立藏语新有效词知识库，(2)利用信息熵对藏语新有效词进行过滤，(3)利用向量空间模块相似度计算提取藏语新有效词。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文本刊更多论文

Automatic Extraction Method of Tibetan New Valid Words

This paper proposes a model to automatically extract Tibetan new valid words. Through building the dynamic Tibetan corpus from 2009 to 2012, which covers more than 18 Tibetan network media of Tibet, Qinghai, Sichuan, Gansu and Yunnan, we research on the key techniques of Tibetan new valid word extraction: (1) using statistical method to establish Tibetan new word knowledge base, (2) using information entropy to filter Tibetan new valid words, (3) using vector space module similarity calculation to extract Tibetan new valid word.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

2012 Fifth International Conference on Intelligent Networks and Intelligent Systems

自引率

0.00%

发文量