基于监督学习的非正式文本勒索软件实体分类

Nurfadilah Ariffini, A. Zainal, M. A. Maarof, Mohamad Nizam Kassim
{"title":"基于监督学习的非正式文本勒索软件实体分类","authors":"Nurfadilah Ariffini, A. Zainal, M. A. Maarof, Mohamad Nizam Kassim","doi":"10.1109/ICoCSec47621.2019.8970801","DOIUrl":null,"url":null,"abstract":"Analyzing text especially in Malware domain is quite challenging. Even with Natural Language Processing approach, it limits with the absence of specific Named Entity recognizer to extract related entities of malware from unstructured data like text. It is essential to automate the process of extracting information such as Ransomware entity from text and the information extracted could be used as knowledge reasoning like profiling the behaviour of Ransomware using the information available on the Internet. Although the text itself is unstructured, informal text like Internet forum has its problems and challenges to perform the analysis and extraction process. Thus, the performance of machine learning in carrying out the classification of entities from this type of text depending on the complexity of the model. Therefore, this paper presents the comparison of few supervised learning techniques (CRF, Naive Bayes, and SVM) for model training in extracting Ransomware entities from unstructured text in terms of their performance.","PeriodicalId":272402,"journal":{"name":"2019 International Conference on Cybersecurity (ICoCSec)","volume":"68 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2019-09-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"2","resultStr":"{\"title\":\"Ransomware Entities Classification with Supervised Learning for Informal Text\",\"authors\":\"Nurfadilah Ariffini, A. Zainal, M. A. Maarof, Mohamad Nizam Kassim\",\"doi\":\"10.1109/ICoCSec47621.2019.8970801\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"Analyzing text especially in Malware domain is quite challenging. Even with Natural Language Processing approach, it limits with the absence of specific Named Entity recognizer to extract related entities of malware from unstructured data like text. It is essential to automate the process of extracting information such as Ransomware entity from text and the information extracted could be used as knowledge reasoning like profiling the behaviour of Ransomware using the information available on the Internet. Although the text itself is unstructured, informal text like Internet forum has its problems and challenges to perform the analysis and extraction process. Thus, the performance of machine learning in carrying out the classification of entities from this type of text depending on the complexity of the model. Therefore, this paper presents the comparison of few supervised learning techniques (CRF, Naive Bayes, and SVM) for model training in extracting Ransomware entities from unstructured text in terms of their performance.\",\"PeriodicalId\":272402,\"journal\":{\"name\":\"2019 International Conference on Cybersecurity (ICoCSec)\",\"volume\":\"68 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2019-09-01\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"2\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"2019 International Conference on Cybersecurity (ICoCSec)\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1109/ICoCSec47621.2019.8970801\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"2019 International Conference on Cybersecurity (ICoCSec)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/ICoCSec47621.2019.8970801","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
引用次数: 2

摘要

文本分析是一项非常具有挑战性的工作,尤其是在恶意软件领域。即使使用自然语言处理方法,由于缺乏特定的命名实体识别器,因此无法从文本等非结构化数据中提取相关的恶意软件实体。从文本中提取勒索软件实体等信息的自动化过程是至关重要的,提取的信息可以用作知识推理,如利用互联网上可用的信息来分析勒索软件的行为。虽然文本本身是非结构化的,但像互联网论坛这样的非正式文本在进行分析和提取过程中存在着问题和挑战。因此,机器学习在从这种类型的文本中执行实体分类时的性能取决于模型的复杂性。因此,本文比较了几种监督学习技术(CRF、朴素贝叶斯和支持向量机)在从非结构化文本中提取勒索软件实体的模型训练中的性能。
本文章由计算机程序翻译,如有差异,请以英文原文为准。
Ransomware Entities Classification with Supervised Learning for Informal Text
Analyzing text especially in Malware domain is quite challenging. Even with Natural Language Processing approach, it limits with the absence of specific Named Entity recognizer to extract related entities of malware from unstructured data like text. It is essential to automate the process of extracting information such as Ransomware entity from text and the information extracted could be used as knowledge reasoning like profiling the behaviour of Ransomware using the information available on the Internet. Although the text itself is unstructured, informal text like Internet forum has its problems and challenges to perform the analysis and extraction process. Thus, the performance of machine learning in carrying out the classification of entities from this type of text depending on the complexity of the model. Therefore, this paper presents the comparison of few supervised learning techniques (CRF, Naive Bayes, and SVM) for model training in extracting Ransomware entities from unstructured text in terms of their performance.
求助全文
通过发布文献求助,成功后即可免费获取论文全文。 去求助
来源期刊
自引率
0.00%
发文量
0
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
确定
请完成安全验证×
copy
已复制链接
快去分享给好友吧!
我知道了
右上角分享
点击右上角分享
0
联系我们:info@booksci.cn Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。 Copyright © 2023 布克学术 All rights reserved.
京ICP备2023020795号-1
ghs 京公网安备 11010802042870号
Book学术文献互助
Book学术文献互助群
群 号:481959085
Book学术官方微信