{"title":"基于产品协同过滤的大型电子商务推荐系统","authors":"Trang Trinh , Van-Ho Nguyen , Nghia Nguyen , Duy-Nghia Nguyen","doi":"10.1016/j.jjimei.2025.100322","DOIUrl":null,"url":null,"abstract":"<div><div>The rapid growth in e-commerce and the increasing diversity of customer preferences necessitates the development of an effective recommender system for a business offering a wide range of products. This paper introduces a product-based collaborative filtering approach utilizing Apache Spark, a powerful parallel processing framework to address the scalability issues of recommender systems in the cloud computing environment. Using Spark's distributed computing ability, our model attains a surprising 7.6 times speedup on the training time compared to traditional single-machine methods while preserving accuracy with a Root Mean Square Error (RMSE) 0.9. These results demonstrate the effectiveness of parallel and distributed techniques in developing efficient and accurate recommender systems for large-scale e-commerce applications. Future work will focus on applying multi-model to enhance the accuracy of prediction and configuration to optimize the cost of cluster operations.</div></div>","PeriodicalId":100699,"journal":{"name":"International Journal of Information Management Data Insights","volume":"5 1","pages":"Article 100322"},"PeriodicalIF":0.0000,"publicationDate":"2025-01-29","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":"{\"title\":\"Product collaborative filtering based recommendation systems for large-scale E-commerce\",\"authors\":\"Trang Trinh , Van-Ho Nguyen , Nghia Nguyen , Duy-Nghia Nguyen\",\"doi\":\"10.1016/j.jjimei.2025.100322\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"<div><div>The rapid growth in e-commerce and the increasing diversity of customer preferences necessitates the development of an effective recommender system for a business offering a wide range of products. This paper introduces a product-based collaborative filtering approach utilizing Apache Spark, a powerful parallel processing framework to address the scalability issues of recommender systems in the cloud computing environment. Using Spark's distributed computing ability, our model attains a surprising 7.6 times speedup on the training time compared to traditional single-machine methods while preserving accuracy with a Root Mean Square Error (RMSE) 0.9. These results demonstrate the effectiveness of parallel and distributed techniques in developing efficient and accurate recommender systems for large-scale e-commerce applications. Future work will focus on applying multi-model to enhance the accuracy of prediction and configuration to optimize the cost of cluster operations.</div></div>\",\"PeriodicalId\":100699,\"journal\":{\"name\":\"International Journal of Information Management Data Insights\",\"volume\":\"5 1\",\"pages\":\"Article 100322\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2025-01-29\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"International Journal of Information Management Data Insights\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://www.sciencedirect.com/science/article/pii/S2667096825000047\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"International Journal of Information Management Data Insights","FirstCategoryId":"1085","ListUrlMain":"https://www.sciencedirect.com/science/article/pii/S2667096825000047","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
Product collaborative filtering based recommendation systems for large-scale E-commerce
The rapid growth in e-commerce and the increasing diversity of customer preferences necessitates the development of an effective recommender system for a business offering a wide range of products. This paper introduces a product-based collaborative filtering approach utilizing Apache Spark, a powerful parallel processing framework to address the scalability issues of recommender systems in the cloud computing environment. Using Spark's distributed computing ability, our model attains a surprising 7.6 times speedup on the training time compared to traditional single-machine methods while preserving accuracy with a Root Mean Square Error (RMSE) 0.9. These results demonstrate the effectiveness of parallel and distributed techniques in developing efficient and accurate recommender systems for large-scale e-commerce applications. Future work will focus on applying multi-model to enhance the accuracy of prediction and configuration to optimize the cost of cluster operations.