Automatic Construction of English-Vietnamese Parallel Corpus through Web Mining

2007 IEEE International Conference on Research, Innovation and Vision for the Future Pub Date : 2007-03-05 DOI:10.1109/RIVF.2007.369166

V. B. Dang, Bao-Quoc Ho

引用次数: 14

Abstract

Parallel corpus has become a very essential resource for multilingual natural language processing and there are large scale of parallel texts available on the Internet these days. In this paper, we propose a simple but reliable method to construct an English-Vietnamese parallel corpus through Web mining. Our system can automatically download and detect parallel Web pages on a given domain to construct a parallel corpus that is well-aligned at paragraph level with completely clean texts. The proposed technique can be easily applied to other language pairs. Experiments have been made and shown promising results.

查看原文本刊更多论文

基于Web挖掘的英越平行语料库自动构建

并行语料库已成为多语种自然语言处理的重要资源，目前互联网上存在大量的并行文本。本文提出了一种简单可靠的基于Web挖掘的英越平行语料库构建方法。我们的系统可以自动下载和检测给定域上的并行网页，以构建一个在段落级别上对齐良好、文本完全干净的并行语料库。所提出的技术可以很容易地应用于其他语言对。已经进行了实验，并显示出令人满意的结果。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

2007 IEEE International Conference on Research, Innovation and Vision for the Future

自引率

0.00%

发文量