{"title":"TK-SMF: Top-K Semantic Mask Fusion With Dynamic Budget and Recovery Against Adversarial Patch Attacks Under Semantic Fragmentation in Vision Models","authors":"İbrahim Yaylalı, Alper Kılıç","doi":"10.1002/cpe.70908","DOIUrl":null,"url":null,"abstract":"<div>\n \n <p>Deep learning-based vision models have severe security vulnerabilities against physical adversarial patch attacks. Previous studies propose explainability-based defense mechanisms (e.g., PatchOut and SALIUITL). These methods aim to recover the system by masking the regions that affect model decisions the most. However, these approaches generally assume that the adversarial effect is located in a single, continuous region (single mask selection). In this study, we empirically demonstrate the “Semantic Fragmentation” and saliency bleeding problems caused by multiple scattered patches in explainability maps. To overcome this issue, we propose the “Top-K Semantic Mask Fusion with Dynamic Budget (TK-SMF)” system. The proposed framework initially extracts explainability maps using LayerGradCAM, subsequently constructs a pool of candidate semantic segments via the Fast Segment Anything Model (FastSAM), and ultimately fuses the Top-K highest-risk masks within a strictly fixed area budget. This budget is anchored to a predefined worst-case threat assumption, eliminating the need to know the attacker's actual patch count during inference. We then restore the selected masked areas using OpenCV Inpainting to achieve model recovery. We evaluated the proposed framework on 38 distinct and challenging physical-world scenarios, encompassing diverse tactical vehicle categories under single- and multipatch attacks. Experimental results prove that TK-SMF breaks the targeted attack illusion with a 100% success rate. The results show that dynamic and flexible defense architectures exceed the limits of the single-mask assumption. They significantly improve model accuracy and reliability against complex physical world threats.</p>\n </div>","PeriodicalId":55214,"journal":{"name":"Concurrency and Computation-Practice & Experience","volume":"38 17","pages":""},"PeriodicalIF":2.2000,"publicationDate":"2026-08-28","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Concurrency and Computation-Practice & Experience","FirstCategoryId":"94","ListUrlMain":"https://onlinelibrary.wiley.com/doi/10.1002/cpe.70908","RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q3","JCRName":"COMPUTER SCIENCE, SOFTWARE ENGINEERING","Score":null,"Total":0}
引用次数: 0
Abstract
Deep learning-based vision models have severe security vulnerabilities against physical adversarial patch attacks. Previous studies propose explainability-based defense mechanisms (e.g., PatchOut and SALIUITL). These methods aim to recover the system by masking the regions that affect model decisions the most. However, these approaches generally assume that the adversarial effect is located in a single, continuous region (single mask selection). In this study, we empirically demonstrate the “Semantic Fragmentation” and saliency bleeding problems caused by multiple scattered patches in explainability maps. To overcome this issue, we propose the “Top-K Semantic Mask Fusion with Dynamic Budget (TK-SMF)” system. The proposed framework initially extracts explainability maps using LayerGradCAM, subsequently constructs a pool of candidate semantic segments via the Fast Segment Anything Model (FastSAM), and ultimately fuses the Top-K highest-risk masks within a strictly fixed area budget. This budget is anchored to a predefined worst-case threat assumption, eliminating the need to know the attacker's actual patch count during inference. We then restore the selected masked areas using OpenCV Inpainting to achieve model recovery. We evaluated the proposed framework on 38 distinct and challenging physical-world scenarios, encompassing diverse tactical vehicle categories under single- and multipatch attacks. Experimental results prove that TK-SMF breaks the targeted attack illusion with a 100% success rate. The results show that dynamic and flexible defense architectures exceed the limits of the single-mask assumption. They significantly improve model accuracy and reliability against complex physical world threats.
期刊介绍:
Concurrency and Computation: Practice and Experience (CCPE) publishes high-quality, original research papers, and authoritative research review papers, in the overlapping fields of:
Parallel and distributed computing;
High-performance computing;
Computational and data science;
Artificial intelligence and machine learning;
Big data applications, algorithms, and systems;
Network science;
Ontologies and semantics;
Security and privacy;
Cloud/edge/fog computing;
Green computing; and
Quantum computing.