对话中话语标记的自动识别:以相似为例

SIGDIAL Workshop Pub Date : 1900-01-01 DOI:10.7892/BORIS.78686

S. Zufferey, Andrei Popescu-Belis

{"title":"对话中话语标记的自动识别:以相似为例","authors":"S. Zufferey, Andrei Popescu-Belis","doi":"10.7892/BORIS.78686","DOIUrl":null,"url":null,"abstract":"This article discusses the detection of discourse \n markers (DM) in dialog transcriptions, \nby human annotators and by automated \nmeans. After a theoretical discussion of the \ndefinition of DMs and their relevance to natural \n language processing, we focus on the role \nof like as a DM. Results from experiments \nwith human annotators show that detection of \nDMs is a difficult but reliable task, which requires \n prosodic information from soundtracks. \nThen, several types of features are defined for \nautomatic disambiguation of like: collocations, \n part-of-speech tags and duration-based \nfeatures. Decision-tree learning shows that for \nlike, nearly 70% precision can be reached, \nwith near 100% recall, mainly using collocation \n filters. Similar results hold for well, with \nabout 91% precision at 100% recall.","PeriodicalId":426429,"journal":{"name":"SIGDIAL Workshop","volume":"10 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"1900-01-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"31","resultStr":"{\"title\":\"Towards Automatic Identification of Discourse Markers in Dialogs: The Case of Like\",\"authors\":\"S. Zufferey, Andrei Popescu-Belis\",\"doi\":\"10.7892/BORIS.78686\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"This article discusses the detection of discourse \\n markers (DM) in dialog transcriptions, \\nby human annotators and by automated \\nmeans. After a theoretical discussion of the \\ndefinition of DMs and their relevance to natural \\n language processing, we focus on the role \\nof like as a DM. Results from experiments \\nwith human annotators show that detection of \\nDMs is a difficult but reliable task, which requires \\n prosodic information from soundtracks. \\nThen, several types of features are defined for \\nautomatic disambiguation of like: collocations, \\n part-of-speech tags and duration-based \\nfeatures. Decision-tree learning shows that for \\nlike, nearly 70% precision can be reached, \\nwith near 100% recall, mainly using collocation \\n filters. Similar results hold for well, with \\nabout 91% precision at 100% recall.\",\"PeriodicalId\":426429,\"journal\":{\"name\":\"SIGDIAL Workshop\",\"volume\":\"10 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"1900-01-01\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"31\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"SIGDIAL Workshop\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.7892/BORIS.78686\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"SIGDIAL Workshop","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.7892/BORIS.78686","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 31

摘要

本文讨论了对话文本中话语标记(DM)的检测方法，包括人工注释器和自动化方法。在对DM的定义及其与自然语言处理的相关性进行了理论讨论之后，我们将重点放在like作为DM的作用上。人类注释器的实验结果表明，检测DM是一项困难但可靠的任务，这需要来自音轨的韵律信息。然后，定义了几种类型的自动消歧特征:搭配、词性标签和基于持续时间的特征。决策树学习表明，对于like，主要使用搭配过滤器，准确率接近70%，召回率接近100%。类似的结果也很好，在100%召回率下，准确率约为91%。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文本刊更多论文

Towards Automatic Identification of Discourse Markers in Dialogs: The Case of Like

This article discusses the detection of discourse markers (DM) in dialog transcriptions, by human annotators and by automated means. After a theoretical discussion of the definition of DMs and their relevance to natural language processing, we focus on the role of like as a DM. Results from experiments with human annotators show that detection of DMs is a difficult but reliable task, which requires prosodic information from soundtracks. Then, several types of features are defined for automatic disambiguation of like: collocations, part-of-speech tags and duration-based features. Decision-tree learning shows that for like, nearly 70% precision can be reached, with near 100% recall, mainly using collocation filters. Similar results hold for well, with about 91% precision at 100% recall.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

SIGDIAL Workshop

自引率

0.00%

发文量