土耳其语文本分类的预训练神经模型

2021 6th International Conference on Computer Science and Engineering (UBMK) Pub Date : 2021-09-15 DOI:10.1109/UBMK52708.2021.9558878

Halil Ibrahim Okur, A. Sertbas

{"title":"土耳其语文本分类的预训练神经模型","authors":"Halil Ibrahim Okur, A. Sertbas","doi":"10.1109/UBMK52708.2021.9558878","DOIUrl":null,"url":null,"abstract":"In the text classification process, which is a sub-task of NLP, the preprocessing and indexing of the text has a direct determining effect on the performance for NLP models. When the studies on pre-trained models are examined, it is seen that the changes made on the models developed for world languages or training the same model with a Turkish text dataset. Word-embedding is considered to be the most critical point of the text processing problem. The two most popular word embedding methods today are Word2Vec and Glove, which embed words into a corpus using multidimensional vectors. BERT, Electra and Fastext models, which have a contextual word representation method and a deep neural network architecture, have been frequently used in the creation of pre-trained models recently. In this study, the use and performance results of pre-trained models on TTC-3600 and TRT-Haber text sets prepared for Turkish text classification NLP task are shown. By using pre-trained models obtained with large corpus, a certain time and hardware cost, the text classification process is performed with less effort and high performance.","PeriodicalId":106516,"journal":{"name":"2021 6th International Conference on Computer Science and Engineering (UBMK)","volume":"16 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2021-09-15","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":"{\"title\":\"Pretrained Neural Models for Turkish Text Classification\",\"authors\":\"Halil Ibrahim Okur, A. Sertbas\",\"doi\":\"10.1109/UBMK52708.2021.9558878\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"In the text classification process, which is a sub-task of NLP, the preprocessing and indexing of the text has a direct determining effect on the performance for NLP models. When the studies on pre-trained models are examined, it is seen that the changes made on the models developed for world languages or training the same model with a Turkish text dataset. Word-embedding is considered to be the most critical point of the text processing problem. The two most popular word embedding methods today are Word2Vec and Glove, which embed words into a corpus using multidimensional vectors. BERT, Electra and Fastext models, which have a contextual word representation method and a deep neural network architecture, have been frequently used in the creation of pre-trained models recently. In this study, the use and performance results of pre-trained models on TTC-3600 and TRT-Haber text sets prepared for Turkish text classification NLP task are shown. By using pre-trained models obtained with large corpus, a certain time and hardware cost, the text classification process is performed with less effort and high performance.\",\"PeriodicalId\":106516,\"journal\":{\"name\":\"2021 6th International Conference on Computer Science and Engineering (UBMK)\",\"volume\":\"16 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2021-09-15\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"2021 6th International Conference on Computer Science and Engineering (UBMK)\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1109/UBMK52708.2021.9558878\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"2021 6th International Conference on Computer Science and Engineering (UBMK)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/UBMK52708.2021.9558878","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 0

摘要

文本分类是自然语言处理的一个子任务，在文本分类过程中，文本的预处理和索引对自然语言处理模型的性能有直接的决定作用。当对预训练模型的研究进行检查时，可以看到对为世界语言开发的模型或使用土耳其文本数据集训练相同模型所做的更改。词嵌入被认为是文本处理中最关键的问题。目前最流行的两种词嵌入方法是Word2Vec和Glove，它们使用多维向量将词嵌入到语料库中。BERT、Electra和Fastext模型具有上下文词表示方法和深度神经网络架构，近年来被广泛用于预训练模型的创建。在本研究中，展示了预训练模型在为土耳其文本分类NLP任务准备的TTC-3600和TRT-Haber文本集上的使用和性能结果。通过使用大量语料库、一定的时间和硬件成本获得的预训练模型，实现了省力、高性能的文本分类过程。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文本刊更多论文

Pretrained Neural Models for Turkish Text Classification

In the text classification process, which is a sub-task of NLP, the preprocessing and indexing of the text has a direct determining effect on the performance for NLP models. When the studies on pre-trained models are examined, it is seen that the changes made on the models developed for world languages or training the same model with a Turkish text dataset. Word-embedding is considered to be the most critical point of the text processing problem. The two most popular word embedding methods today are Word2Vec and Glove, which embed words into a corpus using multidimensional vectors. BERT, Electra and Fastext models, which have a contextual word representation method and a deep neural network architecture, have been frequently used in the creation of pre-trained models recently. In this study, the use and performance results of pre-trained models on TTC-3600 and TRT-Haber text sets prepared for Turkish text classification NLP task are shown. By using pre-trained models obtained with large corpus, a certain time and hardware cost, the text classification process is performed with less effort and high performance.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

2021 6th International Conference on Computer Science and Engineering (UBMK)

自引率

0.00%

发文量