A character elimination algorithm for lossless data compression

Proceedings DCC 2002. Data Compression Conference Pub Date : 2002-04-02 DOI:10.1109/DCC.2002.1000000

Mark Hosang

引用次数: 1

Abstract

Summary form only given. We present a detailed description of a lossless compression algorithm intended for use on files with non-uniform character distributions. This algorithm takes advantage of the relatively small distances between character occurrences once we remove the less frequent characters. This allows it to create a compressed version of the file that, when decompressed, is an exact copy of the file that was compressed. We begin by performing a Burrows-Wheeler (1994) Transform (BWT) on the file. The algorithm scans this BWT file to create a character frequency model for the compression phase. To deal with the issue of bit encoding, we write every number as a byte or sequence of bytes to the compressed file and run an arithmetic encoder after the file has been compiled.

查看原文本刊更多论文

一种用于无损数据压缩的字符消除算法

只提供摘要形式。我们提出了一种无损压缩算法的详细描述，用于具有非均匀字符分布的文件。该算法利用了字符出现之间相对较小的距离，一旦我们删除了不太频繁的字符。这允许它创建文件的压缩版本，当解压缩时，它是被压缩文件的精确副本。我们首先对文件执行Burrows-Wheeler(1994)变换(BWT)。该算法扫描该BWT文件，为压缩阶段创建字符频率模型。为了处理位编码的问题，我们将每个数字作为字节或字节序列写入压缩文件，并在文件编译后运行算术编码器。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

Proceedings DCC 2002. Data Compression Conference

自引率

0.00%

发文量