{"title":"COMFRE:在语言任务中比较词频的可视化工具","authors":"Shane Sheehan, M. Masoodian, S. Luz","doi":"10.1145/3206505.3206547","DOIUrl":null,"url":null,"abstract":"Comparing frequency distributions is a basic task in statistics and in disciplines that rely on statistical analysis, such as corpus linguistics. However, support for comparing word frequencies between different corpora in corpus linguistics tasks such as lexical analysis and corpus-based translation studies, is often limited to fairly basic techniques like tabular word lists. While other visualizations such as word clouds do exist, they are not widely used in linguistic analysis tasks due to their lack of precision, unsuitability to dealing with the high frequencies of common words, and lack of effective mechanisms for direct manipulation. In this paper, we propose a visualization for comparing word frequencies across two corpora using a combination of slope charts and histogram contours. An interactive implementation of this visualization is also presented. The design of visualization, and the development of the prototype, have been guided through the involvement of expert linguist users.","PeriodicalId":330748,"journal":{"name":"Proceedings of the 2018 International Conference on Advanced Visual Interfaces","volume":"106 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2018-05-29","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"5","resultStr":"{\"title\":\"COMFRE: a visualization for comparing word frequencies in linguistic tasks\",\"authors\":\"Shane Sheehan, M. Masoodian, S. Luz\",\"doi\":\"10.1145/3206505.3206547\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"Comparing frequency distributions is a basic task in statistics and in disciplines that rely on statistical analysis, such as corpus linguistics. However, support for comparing word frequencies between different corpora in corpus linguistics tasks such as lexical analysis and corpus-based translation studies, is often limited to fairly basic techniques like tabular word lists. While other visualizations such as word clouds do exist, they are not widely used in linguistic analysis tasks due to their lack of precision, unsuitability to dealing with the high frequencies of common words, and lack of effective mechanisms for direct manipulation. In this paper, we propose a visualization for comparing word frequencies across two corpora using a combination of slope charts and histogram contours. An interactive implementation of this visualization is also presented. The design of visualization, and the development of the prototype, have been guided through the involvement of expert linguist users.\",\"PeriodicalId\":330748,\"journal\":{\"name\":\"Proceedings of the 2018 International Conference on Advanced Visual Interfaces\",\"volume\":\"106 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2018-05-29\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"5\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"Proceedings of the 2018 International Conference on Advanced Visual Interfaces\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1145/3206505.3206547\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"Proceedings of the 2018 International Conference on Advanced Visual Interfaces","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1145/3206505.3206547","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
COMFRE: a visualization for comparing word frequencies in linguistic tasks
Comparing frequency distributions is a basic task in statistics and in disciplines that rely on statistical analysis, such as corpus linguistics. However, support for comparing word frequencies between different corpora in corpus linguistics tasks such as lexical analysis and corpus-based translation studies, is often limited to fairly basic techniques like tabular word lists. While other visualizations such as word clouds do exist, they are not widely used in linguistic analysis tasks due to their lack of precision, unsuitability to dealing with the high frequencies of common words, and lack of effective mechanisms for direct manipulation. In this paper, we propose a visualization for comparing word frequencies across two corpora using a combination of slope charts and histogram contours. An interactive implementation of this visualization is also presented. The design of visualization, and the development of the prototype, have been guided through the involvement of expert linguist users.