Lihua Yang , Yang Xiao , Zhipeng Tan , Fang Wang , Weizhao Lin , Wei Zhang , Jiaxin Li , Kai Lu
{"title":"一种轻量级的细粒度方案,用于区分热数据的热度,以减少段清理开销","authors":"Lihua Yang , Yang Xiao , Zhipeng Tan , Fang Wang , Weizhao Lin , Wei Zhang , Jiaxin Li , Kai Lu","doi":"10.1016/j.jpdc.2025.105183","DOIUrl":null,"url":null,"abstract":"<div><div>With the widespread adoption of flash memory, the Flash Friendly File System (F2FS) designed to flash memory characteristics has become widely-used in large data centers. However, F2FS encounters from significant cleaning overheads due to its logging scheme writes. We observe that warm data in F2FS account for a substantial proportion, at least 80 %. Nevertheless, the mixed storage of warm data with varying hotness exacerbates segment cleaning challenges. To address this issue, we propose a scheme called M2H, which involves a fine-grained management of warm data hotness identified by the K-means clustering algorithm. M2H determines hotness by considering factors such as file block update distance, most recently used distance, and workload characteristics. M2H facilitates <strong>M</strong>ulti-log delayed writing and <strong>M</strong>odified segment cleaning based on <strong>H</strong>otness. To reduce costs associated with distinguishing data hotness at the file block level, we employ Mini Batch K-means, which is referred to as HMBK. Moreover, for servers equipped with GPUs, the clustering process can be offloaded to the GPU, known as HGPU. We conduct a comprehensive comparison of traditional F2FS, M2H, HMBK, and HGPU on a real platform. Results show that compared to traditional F2FS, HGPU reduces the number of segment cleanings by 54.41 % to 97.93 %.</div></div>","PeriodicalId":54775,"journal":{"name":"Journal of Parallel and Distributed Computing","volume":"207 ","pages":"Article 105183"},"PeriodicalIF":3.4000,"publicationDate":"2026-01-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":"{\"title\":\"A lightweight fine-grained scheme for distinguishing the hotness of warm data to reduce segment cleaning overhead\",\"authors\":\"Lihua Yang , Yang Xiao , Zhipeng Tan , Fang Wang , Weizhao Lin , Wei Zhang , Jiaxin Li , Kai Lu\",\"doi\":\"10.1016/j.jpdc.2025.105183\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"<div><div>With the widespread adoption of flash memory, the Flash Friendly File System (F2FS) designed to flash memory characteristics has become widely-used in large data centers. However, F2FS encounters from significant cleaning overheads due to its logging scheme writes. We observe that warm data in F2FS account for a substantial proportion, at least 80 %. Nevertheless, the mixed storage of warm data with varying hotness exacerbates segment cleaning challenges. To address this issue, we propose a scheme called M2H, which involves a fine-grained management of warm data hotness identified by the K-means clustering algorithm. M2H determines hotness by considering factors such as file block update distance, most recently used distance, and workload characteristics. M2H facilitates <strong>M</strong>ulti-log delayed writing and <strong>M</strong>odified segment cleaning based on <strong>H</strong>otness. To reduce costs associated with distinguishing data hotness at the file block level, we employ Mini Batch K-means, which is referred to as HMBK. Moreover, for servers equipped with GPUs, the clustering process can be offloaded to the GPU, known as HGPU. We conduct a comprehensive comparison of traditional F2FS, M2H, HMBK, and HGPU on a real platform. Results show that compared to traditional F2FS, HGPU reduces the number of segment cleanings by 54.41 % to 97.93 %.</div></div>\",\"PeriodicalId\":54775,\"journal\":{\"name\":\"Journal of Parallel and Distributed Computing\",\"volume\":\"207 \",\"pages\":\"Article 105183\"},\"PeriodicalIF\":3.4000,\"publicationDate\":\"2026-01-01\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"Journal of Parallel and Distributed Computing\",\"FirstCategoryId\":\"94\",\"ListUrlMain\":\"https://www.sciencedirect.com/science/article/pii/S0743731525001509\",\"RegionNum\":3,\"RegionCategory\":\"计算机科学\",\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"2025/10/10 0:00:00\",\"PubModel\":\"Epub\",\"JCR\":\"Q1\",\"JCRName\":\"COMPUTER SCIENCE, THEORY & METHODS\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"Journal of Parallel and Distributed Computing","FirstCategoryId":"94","ListUrlMain":"https://www.sciencedirect.com/science/article/pii/S0743731525001509","RegionNum":3,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"2025/10/10 0:00:00","PubModel":"Epub","JCR":"Q1","JCRName":"COMPUTER SCIENCE, THEORY & METHODS","Score":null,"Total":0}
A lightweight fine-grained scheme for distinguishing the hotness of warm data to reduce segment cleaning overhead
With the widespread adoption of flash memory, the Flash Friendly File System (F2FS) designed to flash memory characteristics has become widely-used in large data centers. However, F2FS encounters from significant cleaning overheads due to its logging scheme writes. We observe that warm data in F2FS account for a substantial proportion, at least 80 %. Nevertheless, the mixed storage of warm data with varying hotness exacerbates segment cleaning challenges. To address this issue, we propose a scheme called M2H, which involves a fine-grained management of warm data hotness identified by the K-means clustering algorithm. M2H determines hotness by considering factors such as file block update distance, most recently used distance, and workload characteristics. M2H facilitates Multi-log delayed writing and Modified segment cleaning based on Hotness. To reduce costs associated with distinguishing data hotness at the file block level, we employ Mini Batch K-means, which is referred to as HMBK. Moreover, for servers equipped with GPUs, the clustering process can be offloaded to the GPU, known as HGPU. We conduct a comprehensive comparison of traditional F2FS, M2H, HMBK, and HGPU on a real platform. Results show that compared to traditional F2FS, HGPU reduces the number of segment cleanings by 54.41 % to 97.93 %.
期刊介绍:
This international journal is directed to researchers, engineers, educators, managers, programmers, and users of computers who have particular interests in parallel processing and/or distributed computing.
The Journal of Parallel and Distributed Computing publishes original research papers and timely review articles on the theory, design, evaluation, and use of parallel and/or distributed computing systems. The journal also features special issues on these topics; again covering the full range from the design to the use of our targeted systems.