On optimality of data clustering for packet-level memory-assisted compression of network traffic

2014 IEEE 15th International Workshop on Signal Processing Advances in Wireless Communications (SPAWC) Pub Date : 2014-06-22 DOI:10.1109/SPAWC.2014.6941726

Ahmad Beirami, Liling Huang, Mohsen Sardari, F. Fekri

{"title":"On optimality of data clustering for packet-level memory-assisted compression of network traffic","authors":"Ahmad Beirami, Liling Huang, Mohsen Sardari, F. Fekri","doi":"10.1109/SPAWC.2014.6941726","DOIUrl":null,"url":null,"abstract":"Recently, we proposed a framework called memory-assisted compression that learns the statistical properties of the sequence-generating server at intermediate network nodes and then leverages the learnt models to overcome the inevitable redundancy (overhead) in the universal compression of the payloads of the short-length network packets. In this paper, we prove that when the content-generating server is comprised of a mixture of parametric sources, label-based clustering of the data to their original sequence-generating models from the mixture is optimal almost surely as it achieves the mixture entropy (which is the lower bound on the average codeword length). Motivated by this result, we present a K-means clustering technique as the proof of concept to demonstrate the benefits of memory-assisted compression performance. Simulation results confirm the effectiveness of the proposed approach by matching the expected improvements predicted by theory on man-made mixture sources. Finally, the benefits of the cluster-based memory-assisted compression are validated on real data traffic traces demonstrating more than 50% traffic reduction on average in data gathered from wireless users.","PeriodicalId":420837,"journal":{"name":"2014 IEEE 15th International Workshop on Signal Processing Advances in Wireless Communications (SPAWC)","volume":"12 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2014-06-22","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"2","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"2014 IEEE 15th International Workshop on Signal Processing Advances in Wireless Communications (SPAWC)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/SPAWC.2014.6941726","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 2

Abstract

Recently, we proposed a framework called memory-assisted compression that learns the statistical properties of the sequence-generating server at intermediate network nodes and then leverages the learnt models to overcome the inevitable redundancy (overhead) in the universal compression of the payloads of the short-length network packets. In this paper, we prove that when the content-generating server is comprised of a mixture of parametric sources, label-based clustering of the data to their original sequence-generating models from the mixture is optimal almost surely as it achieves the mixture entropy (which is the lower bound on the average codeword length). Motivated by this result, we present a K-means clustering technique as the proof of concept to demonstrate the benefits of memory-assisted compression performance. Simulation results confirm the effectiveness of the proposed approach by matching the expected improvements predicted by theory on man-made mixture sources. Finally, the benefits of the cluster-based memory-assisted compression are validated on real data traffic traces demonstrating more than 50% traffic reduction on average in data gathered from wireless users.

查看原文本刊更多论文

数据包级内存辅助网络流量压缩中数据聚类的最优性

最近，我们提出了一种称为内存辅助压缩的框架，该框架学习中间网络节点上序列生成服务器的统计属性，然后利用学习到的模型来克服短长度网络数据包有效负载通用压缩中不可避免的冗余(开销)。在本文中，我们证明了当内容生成服务器由参数源的混合物组成时，基于标签的数据聚类到其原始序列生成模型几乎肯定是最优的，因为它实现了混合熵(这是平均码字长度的下界)。受此结果的启发，我们提出了K-means聚类技术作为概念证明，以证明内存辅助压缩性能的好处。仿真结果验证了该方法的有效性，与理论预测的人造混合源的预期改进相吻合。最后，基于集群的内存辅助压缩的好处在实际数据流量跟踪中得到验证，表明从无线用户收集的数据平均减少了50%以上的流量。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

2014 IEEE 15th International Workshop on Signal Processing Advances in Wireless Communications (SPAWC)

自引率

0.00%

发文量