Lucian Radu Teodorescu, Razvan Boldizsar, Mihai Alexandru Ordean, Melania Duma, Laura Detesan, Mihaela Ordean
{"title":"Part of Speech Tagging for Romanian Text-to-Speech System","authors":"Lucian Radu Teodorescu, Razvan Boldizsar, Mihai Alexandru Ordean, Melania Duma, Laura Detesan, Mihaela Ordean","doi":"10.1109/SYNASC.2011.55","DOIUrl":null,"url":null,"abstract":"This paper describes a Part of Speech (POS) tagger that has been developed for Romanian Text-to-Speech purposes. In our Text-to-Speech (TTS) system, the Part of Speech tagger is used to disambiguate the pronunciation of some homograph words, determine the semantic links between words, phrase breaks and intonation phrase boundaries and eventually design the intonation curves. The paper focuses on the development and evaluation of the Romanian POS tagger. The findings of this paper show that Naive Bayes models can very well be used for tagging in a hybrid system composed of trained statistical model and a word database. Our experimental results have uncovered an acceptable accuracy and real time performance of the integrated model using a reduced tag set.","PeriodicalId":184344,"journal":{"name":"2011 13th International Symposium on Symbolic and Numeric Algorithms for Scientific Computing","volume":"209 2 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2011-09-26","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"2","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"2011 13th International Symposium on Symbolic and Numeric Algorithms for Scientific Computing","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/SYNASC.2011.55","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
引用次数: 2
Abstract
This paper describes a Part of Speech (POS) tagger that has been developed for Romanian Text-to-Speech purposes. In our Text-to-Speech (TTS) system, the Part of Speech tagger is used to disambiguate the pronunciation of some homograph words, determine the semantic links between words, phrase breaks and intonation phrase boundaries and eventually design the intonation curves. The paper focuses on the development and evaluation of the Romanian POS tagger. The findings of this paper show that Naive Bayes models can very well be used for tagging in a hybrid system composed of trained statistical model and a word database. Our experimental results have uncovered an acceptable accuracy and real time performance of the integrated model using a reduced tag set.