Int. J. Comput. Linguistics Chin. Lang. Process.最新文献_第3页

Modeling Taiwanese POS Tagging Using Statistical Methods and Mandarin Training Data 基于统计方法和普通话训练数据的台湾词性标注建模

Int. J. Comput. Linguistics Chin. Lang. Process. Pub Date : 2009-09-01 DOI: 10.30019/IJCLCLP.200909.0001

Un-Gian Iunn, Jia-hung Tai, K. Lau, Cheng-Yan Kao, Keh-Jiann Chen

引用次数: 1

Automatic Recognition of Cantonese-English Code-Mixing Speech 粤英混码语音的自动识别

Int. J. Comput. Linguistics Chin. Lang. Process. Pub Date : 2009-09-01 DOI: 10.30019/IJCLCLP.200909.0003

Joyce Y. C. Chan, Houwei Cao, P. Ching, Tan Lee

{"title":"Automatic Recognition of Cantonese-English Code-Mixing Speech","authors":"Joyce Y. C. Chan, Houwei Cao, P. Ching, Tan Lee","doi":"10.30019/IJCLCLP.200909.0003","DOIUrl":"https://doi.org/10.30019/IJCLCLP.200909.0003","url":null,"abstract":"Code-mixing is a common phenomenon in bilingual societies. It refers to the intra-sentential switching of two different languages in a spoken utterance. This paper presents the first study on automatic recognition of Cantonese-English code-mixing speech, which is common in Hong Kong. This study starts with the design and compilation of code-mixing speech and text corpora. The problems of acoustic modeling, language modeling, and language boundary detection are investigated. Subsequently, a large-vocabulary code-mixing speech recognition system is developed based on a two-pass decoding algorithm. For acoustic modeling, it is shown that cross-lingual acoustic models are more appropriate than language-dependent models. The language models being used are character tri-grams, in which the embedded English words are grouped into a small number of classes. Language boundary detection is done either by exploiting the phonological and lexical differences between the two languages or is done based on the result of cross-lingual speech recognition. The language boundary information is used to re-score the hypothesized syllables or words in the decoding process. The proposed code-mixing speech recognition system attains the accuracies of 56.4% and 53.0% for the Cantonese syllables and English words in code-mixing utterances.","PeriodicalId":436300,"journal":{"name":"Int. J. Comput. Linguistics Chin. Lang. Process.","volume":"63 1","pages":"0"},"PeriodicalIF":0.0,"publicationDate":"2009-09-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"126347014","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}

引用次数: 37

Fertility-based Source-Language-biased Inversion Transduction Grammar for Word Alignment 基于生育的源语言偏倚倒转转导语法的词对齐

Int. J. Comput. Linguistics Chin. Lang. Process. Pub Date : 2009-03-01 DOI: 10.30019/IJCLCLP.200903.0001

Chung-Chi Huang, Jason J. S. Chang

引用次数: 0

Assessing Text Readability Using Hierarchical Lexical Relations Retrieved from WordNet 利用从WordNet检索的分层词法关系评估文本可读性

Int. J. Comput. Linguistics Chin. Lang. Process. Pub Date : 2009-03-01 DOI: 10.30019/IJCLCLP.200903.0003

Shu-Yen Lin, Cheng-chao Su, Yuda Lai, Li-Chin Yang, S. Hsieh

{"title":"Assessing Text Readability Using Hierarchical Lexical Relations Retrieved from WordNet","authors":"Shu-Yen Lin, Cheng-chao Su, Yuda Lai, Li-Chin Yang, S. Hsieh","doi":"10.30019/IJCLCLP.200903.0003","DOIUrl":"https://doi.org/10.30019/IJCLCLP.200903.0003","url":null,"abstract":"Although some traditional readability formulas have shown high predictive validity in the r=0.8 range and above (Chall & Dale, 1995), they are generally not based on genuine linguistic processing factors, but on statistical correlations (Crossley et al., 2008). Improvement of readability assessment should focus on finding variables that truly represent the comprehensibility of text as well as the indices that accurately measure the correlations. In this study, we explore the hierarchical relations between lexical items based on the conceptual categories advanced from Prototype Theory (Rosch et al., 1976). According to this theory and its development, basic level words like guitar represent the objects humans interact with most readily. They are acquired by children earlier than their superordinate words like stringed instrument and their subordinate words like acoustic guitar. Accordingly, the readability of a text is presumably associated with the ratio of basic level words it contains. WordNet (Fellbaum, 1998), a network of meaningfully related words, provides the best online open source database for studying such lexical relations. Our study shows that a basic level noun can be identified by its ratio of forming compounds (e.g. chair→armchair) and the length difference in relation to its hyponyms. We compared graded readings for American children and high school English readings for Taiwanese students by several readability formulas and in terms of basic level noun ratios (i.e. the number of basic level noun types divided by the number of noun types in a text). It is suggested that basic level noun ratios provide a robust and meaningful index of lexical complexity, which is directly associated with text readability.","PeriodicalId":436300,"journal":{"name":"Int. J. Comput. Linguistics Chin. Lang. Process.","volume":"12 1","pages":"0"},"PeriodicalIF":0.0,"publicationDate":"2009-03-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"128067627","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}

引用次数: 13

Corpus Cleanup of Mistaken Agreement Using Word Sense Disambiguation 用词义消歧法清理语料库中的一致性错误

Int. J. Comput. Linguistics Chin. Lang. Process. Pub Date : 2008-12-01 DOI: 10.30019/IJCLCLP.200812.0002

Liang-Chih Yu, Chung-Hsien Wu, Jui-Feng Yeh, E. Hovy

{"title":"Corpus Cleanup of Mistaken Agreement Using Word Sense Disambiguation","authors":"Liang-Chih Yu, Chung-Hsien Wu, Jui-Feng Yeh, E. Hovy","doi":"10.30019/IJCLCLP.200812.0002","DOIUrl":"https://doi.org/10.30019/IJCLCLP.200812.0002","url":null,"abstract":"Word sense annotated corpora are useful resources for many text mining applications. Such corpora are only useful if their annotations are consistent. Most large-scale annotation efforts take special measures to reconcile inter-annotator disagreement. To date, however, nobody has investigated how to automatically determine exemplars in which the annotators agree but are wrong. In this paper, we use OntoNotes, a large-scale corpus of semantic annotations, including word senses, predicate-argument structure, ontology linking, and coreference. To determine the mistaken agreements in word sense annotation, we employ word sense disambiguation (WSD) to select a set of suspicious candidates for human evaluation. Experiments are conducted from three aspects (precision, cost-effectiveness ratio, and entropy) to examine the performance of WSD. The experimental results show that WSD is most effective in identifying erroneous annotations for highly-ambiguous words, while a baseline is better for other cases. The two methods can be combined to improve the cleanup process. This procedure allows us to find approximately 2% of the remaining erroneous agreements in the OntoNotes corpus. A similar procedure can be easily defined to check other annotated corpora.","PeriodicalId":436300,"journal":{"name":"Int. J. Comput. Linguistics Chin. Lang. Process.","volume":"27 1","pages":"0"},"PeriodicalIF":0.0,"publicationDate":"2008-12-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"115990115","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}

引用次数: 0

Hierarchical Taxonomy Integration Using Semantic Feature Expansion on Category-Specific Terms 基于特定类别术语语义特征扩展的分层分类法集成

Int. J. Comput. Linguistics Chin. Lang. Process. Pub Date : 2008-12-01 DOI: 10.30019/IJCLCLP.200812.0003

Cheng-Zen Yang, Ing-Xiang Chen, Cheng-Tse Hung, Ping-Jung Wu

引用次数: 1

Automatic Wikibook Prototyping via Mining Wikipedia 通过挖掘维基百科自动创建维基教科书原型

Int. J. Comput. Linguistics Chin. Lang. Process. Pub Date : 2008-12-01 DOI: 10.30019/IJCLCLP.200812.0004

Jen-Liang Chou, Shih-Hung Wu

引用次数: 0

Feature Weighting Random Forest for Detection of Hidden Web Search Interfaces 基于特征加权随机森林的隐藏Web搜索界面检测

Int. J. Comput. Linguistics Chin. Lang. Process. Pub Date : 2008-12-01 DOI: 10.30019/IJCLCLP.200812.0001

Yunming Ye, Hongbo Li, Xiaobai Deng, J. Huang

引用次数: 19

Knowledge Representation and Sense Disambiguation for Interrogatives in E-HowNet E-HowNet中疑问句的知识表示与语义消歧

Int. J. Comput. Linguistics Chin. Lang. Process. Pub Date : 2008-09-01 DOI: 10.30019/IJCLCLP.200809.0001

Shu-Ling Huang, Keh-Jiann Chen

引用次数: 1

Improved Minimum Phone Error based Discriminative Training of Acoustic Models for Mandarin Large Vocabulary Continuous Speech Recognition 基于声学模型判别训练的普通话大词汇量连续语音识别

Int. J. Comput. Linguistics Chin. Lang. Process. Pub Date : 2008-09-01 DOI: 10.30019/IJCLCLP.200809.0005

Shih-Hung Liu, Fang-Hui Chu, Yueng-Tien Lo, Berlin Chen

{"title":"Improved Minimum Phone Error based Discriminative Training of Acoustic Models for Mandarin Large Vocabulary Continuous Speech Recognition","authors":"Shih-Hung Liu, Fang-Hui Chu, Yueng-Tien Lo, Berlin Chen","doi":"10.30019/IJCLCLP.200809.0005","DOIUrl":"https://doi.org/10.30019/IJCLCLP.200809.0005","url":null,"abstract":"This paper considers minimum phone error (MPE) based discriminative training of acoustic models for Mandarin broadcast news recognition. We present a new phone accuracy function based on the frame-level accuracy of hypothesized phone arcs instead of using the raw phone accuracy function of MPE training. Moreover, a novel data selection approach based on the frame-level normalized entropy of Gaussian posterior probabilities obtained from the word lattice of the training utterance is explored. It has the merit of making the training algorithm focus much more on the training statistics of those frame samples that center nearly around the decision boundary for better discrimination. The underlying characteristics of the presented approaches are extensively investigated, and their performance is verified by comparison with the standard MPE training approach as well as the other related work. Experiments conducted on broadcast news collected in Taiwan demonstrate that the integration of the frame-level phone accuracy calculation and data selection yields slight but consistent improvements over the baseline system.","PeriodicalId":436300,"journal":{"name":"Int. J. Comput. Linguistics Chin. Lang. Process.","volume":"305 1","pages":"0"},"PeriodicalIF":0.0,"publicationDate":"2008-09-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"123120802","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}

引用次数: 2