The dataminer's guide to scalable mixed-membership and nonparametric bayesian models

Proceedings of the 19th ACM SIGKDD international conference on Knowledge discovery and data mining Pub Date : 2013-08-11 DOI:10.1145/2487575.2506181

Amr Ahmed, Alex Smola

引用次数: 0

Abstract

Large amounts of data arise in a multitude of situations, ranging from bioinformatics to astronomy, manufacturing, and medical applications. For concreteness our tutorial focuses on data obtained in the context of the internet, such as user generated content (microblogs, e-mails, messages), behavioral data (locations, interactions, clicks, queries), and graphs. Due to its magnitude, much of the challenges are to extract structure and interpretable models without the need for additional labels, i.e. to design effective unsupervised techniques. We present design patterns for hierarchical nonparametric Bayesian models, efficient inference algorithms, and modeling tools to describe salient aspects of the data.

查看原文本刊更多论文

可扩展混合成员和非参数贝叶斯模型的数据挖掘指南

从生物信息学到天文学、制造业和医学应用，在多种情况下都会产生大量数据。具体而言，我们的教程侧重于在互联网上下文中获得的数据，例如用户生成的内容(微博、电子邮件、消息)、行为数据(位置、交互、点击、查询)和图表。由于其规模，许多挑战是在不需要额外标签的情况下提取结构和可解释模型，即设计有效的无监督技术。我们提出了分层非参数贝叶斯模型的设计模式，有效的推理算法和建模工具来描述数据的突出方面。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

Proceedings of the 19th ACM SIGKDD international conference on Knowledge discovery and data mining

自引率

0.00%

发文量