{"title":"An N-Gram Based Approach to Automatically Identifying Web Page Genre","authors":"Jane E. Mason, M. Shepherd, Jack Duffy","doi":"10.1109/HICSS.2009.581","DOIUrl":null,"url":null,"abstract":"The research reported in this paper is the first phase of a larger project on the automatic classification of web pages by their genres, using n-gram representations of the web pages. In this study, the textual content of web pages is used to create feature sets consisting of the most frequent n-grams and their associated frequencies. We present three methods, each of which uses a distance measure to determine the dissimilarity between two feature sets. Each method forms a feature set for every web page in the test set, however the formation of feature sets from the training set differs between methods: we experiment using one feature set per web page, per genre, and a combination of genre-based feature sets supplemented by subgenre feature sets. We present results for a balanced corpus of seven genres (blog, eshop, FAQs, front page, listing, home page, and search page). Initial results are encouraging.","PeriodicalId":211759,"journal":{"name":"2009 42nd Hawaii International Conference on System Sciences","volume":"80 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2009-01-20","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"33","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"2009 42nd Hawaii International Conference on System Sciences","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/HICSS.2009.581","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
引用次数: 33
Abstract
The research reported in this paper is the first phase of a larger project on the automatic classification of web pages by their genres, using n-gram representations of the web pages. In this study, the textual content of web pages is used to create feature sets consisting of the most frequent n-grams and their associated frequencies. We present three methods, each of which uses a distance measure to determine the dissimilarity between two feature sets. Each method forms a feature set for every web page in the test set, however the formation of feature sets from the training set differs between methods: we experiment using one feature set per web page, per genre, and a combination of genre-based feature sets supplemented by subgenre feature sets. We present results for a balanced corpus of seven genres (blog, eshop, FAQs, front page, listing, home page, and search page). Initial results are encouraging.