Tag based models of English text

Proceedings DCC '98 Data Compression Conference (Cat. No.98TB100225) Pub Date : 1998-03-30 DOI:10.1109/DCC.1998.672130

W. Teahan, J. Cleary

引用次数: 15

Abstract

The problem of compressing English text is important both because of the ubiquity of English as a target for compression and because of the light that compression can shed on the structure of English. English text is examined in conjunction with additional information about the parts of speech of each word in the text (these are referred to as "tags"). It is shown that the tags plus the text can be compressed more than the text alone. Essentially the tags can be compressed for nothing or even a small net saving in size. A comparison is made of a number of different ways of integrating compression of tags and text using an escape mechanism similar to PPM. These are also compared with standard word based and character based compression programs. The result is that the tag and word based schemes always outperform the character based schemes. Overall, the tag based schemes outperform the word based schemes. We conclude by conjecturing that tags chosen for compression rather than linguistic purposes would perform even better.

查看原文本刊更多论文

基于标签的英语文本模型

英语文本的压缩问题之所以重要，一方面是因为英语作为压缩对象无处不在，另一方面是因为压缩可以揭示英语的结构。英语文本是与文本中每个单词词性的附加信息(这些被称为“标签”)一起检查的。结果表明，标签加文本比单独压缩文本更有效。基本上，标签可以免费压缩，甚至可以节省少量的净尺寸。对使用类似于PPM的转义机制集成标记和文本压缩的许多不同方法进行了比较。这些还与标准的基于单词和基于字符的压缩程序进行了比较。结果是基于标记和词的方案总是优于基于字符的方案。总的来说，基于标签的方案优于基于词的方案。我们推断，选择用于压缩而不是语言目的的标签会表现得更好。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

Proceedings DCC '98 Data Compression Conference (Cat. No.98TB100225)

自引率

0.00%

发文量