From Few-Shot to Zero-Shot: Towards Generalist Graph Anomaly Detection

IF 11.6 2区 计算机科学 Q1 COMPUTER SCIENCE, ARTIFICIAL INTELLIGENCE
Yixin Liu;Shiyuan Li;Yu Zheng;Qingfeng Chen;Chengqi Zhang;Philip S. Yu;Shirui Pan
{"title":"From Few-Shot to Zero-Shot: Towards Generalist Graph Anomaly Detection","authors":"Yixin Liu;Shiyuan Li;Yu Zheng;Qingfeng Chen;Chengqi Zhang;Philip S. Yu;Shirui Pan","doi":"10.1109/TKDE.2026.3691902","DOIUrl":null,"url":null,"abstract":"Graph anomaly detection (GAD) is critical for identifying abnormal nodes in graph-structured data from diverse domains, including cybersecurity and social networks. The existing GAD methods often focus on the learning paradigms of “one-model-for-one-dataset”, requiring dataset-specific training for each dataset to achieve optimal performance. However, this paradigm suffers from limitations, such as high computational and data costs, limited generalization and transferability to new datasets, and challenges in privacy-sensitive scenarios where access to full datasets or sufficient labels is restricted. To address these limitations, we propose a novel generalist GAD paradigm that aims to develop a unified model capable of detecting anomalies on multiple unseen datasets without retraining/fine-tuning or customization. To this end, we propose a few-shot generalist GAD method with three key designs, namely feature <u>A</u>lignment, a <u>R</u>esidual encoder, and in-<u>C</u>ontext learning, abbreviated as ARC. As a generalist approach, ARC only requires a few labeled normal samples during prediction on any unseen graphs. Specifically, ARC consists of three modules: a feature <u>A</u>lignment module to unify and align features across datasets, a <u>R</u>esidual graph encoder to capture dataset-agnostic anomaly representations, and a cross-attentive in-<u>C</u>ontext learning module to score anomalies using few-shot normal context. Building on ARC, we further introduce ARC<inline-formula><tex-math>$_{\\mathrm{zero}}$</tex-math></inline-formula> for the zero-shot generalist GAD setting, which selects representative pseudo-normal nodes via a pseudo-context mechanism and thus enables fully label-free inference on unseen datasets. Experiments on 17 real-world datasets demonstrate that ARC and ARC<inline-formula><tex-math>$_{\\mathrm{zero}}$</tex-math></inline-formula> effectively detect anomalies, exhibit strong generalization ability, and perform efficiently under few-shot and zero-shot settings.","PeriodicalId":13496,"journal":{"name":"IEEE Transactions on Knowledge and Data Engineering","volume":"38 7","pages":"4357-4372"},"PeriodicalIF":11.6000,"publicationDate":"2026-07-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"IEEE Transactions on Knowledge and Data Engineering","FirstCategoryId":"94","ListUrlMain":"https://ieeexplore.ieee.org/document/11514102/","RegionNum":2,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"2026/3/11 0:00:00","PubModel":"Epub","JCR":"Q1","JCRName":"COMPUTER SCIENCE, ARTIFICIAL INTELLIGENCE","Score":null,"Total":0}
引用次数: 0

Abstract

Graph anomaly detection (GAD) is critical for identifying abnormal nodes in graph-structured data from diverse domains, including cybersecurity and social networks. The existing GAD methods often focus on the learning paradigms of “one-model-for-one-dataset”, requiring dataset-specific training for each dataset to achieve optimal performance. However, this paradigm suffers from limitations, such as high computational and data costs, limited generalization and transferability to new datasets, and challenges in privacy-sensitive scenarios where access to full datasets or sufficient labels is restricted. To address these limitations, we propose a novel generalist GAD paradigm that aims to develop a unified model capable of detecting anomalies on multiple unseen datasets without retraining/fine-tuning or customization. To this end, we propose a few-shot generalist GAD method with three key designs, namely feature Alignment, a Residual encoder, and in-Context learning, abbreviated as ARC. As a generalist approach, ARC only requires a few labeled normal samples during prediction on any unseen graphs. Specifically, ARC consists of three modules: a feature Alignment module to unify and align features across datasets, a Residual graph encoder to capture dataset-agnostic anomaly representations, and a cross-attentive in-Context learning module to score anomalies using few-shot normal context. Building on ARC, we further introduce ARC$_{\mathrm{zero}}$ for the zero-shot generalist GAD setting, which selects representative pseudo-normal nodes via a pseudo-context mechanism and thus enables fully label-free inference on unseen datasets. Experiments on 17 real-world datasets demonstrate that ARC and ARC$_{\mathrm{zero}}$ effectively detect anomalies, exhibit strong generalization ability, and perform efficiently under few-shot and zero-shot settings.
从少射到零射:通才图异常检测
图异常检测(GAD)对于识别来自网络安全和社交网络等不同领域的图结构数据中的异常节点至关重要。现有的GAD方法往往侧重于“一个模型对应一个数据集”的学习范式,需要对每个数据集进行特定数据集的训练以达到最佳性能。然而,这种模式存在局限性,例如高计算和数据成本,有限的泛化和新数据集的可移植性,以及在访问完整数据集或足够标签受到限制的隐私敏感场景中的挑战。为了解决这些限制,我们提出了一种新的通用GAD范式,旨在开发一个统一的模型,能够检测多个未见过的数据集上的异常,而无需重新训练/微调或定制。为此,我们提出了一种具有三个关键设计的少镜头通才GAD方法,即特征对齐,残差编码器和上下文学习(简称ARC)。作为一种通才方法,ARC在预测任何未见过的图时只需要少量标记的正态样本。具体来说,ARC由三个模块组成:一个特征对齐模块,用于统一和对齐跨数据集的特征,一个残差图编码器,用于捕获与数据集无关的异常表示,以及一个交叉关注的上下文学习模块,用于使用少量正常上下文对异常进行评分。在ARC的基础上,我们进一步引入ARC$ {\ maththrm {zero}}$用于零射通才GAD设置,它通过伪上下文机制选择具有代表性的伪正态节点,从而实现对未见数据集的完全无标签推断。在17个真实数据集上的实验表明,ARC和ARC$_{\ maththrm {zero}}$能够有效地检测异常,具有较强的泛化能力,在few-shot和zero-shot设置下都能有效地执行异常。
本文章由计算机程序翻译,如有差异,请以英文原文为准。
求助全文
约1分钟内获得全文 求助全文
来源期刊
IEEE Transactions on Knowledge and Data Engineering
IEEE Transactions on Knowledge and Data Engineering 工程技术-工程:电子与电气
CiteScore
11.70
自引率
3.40%
发文量
515
审稿时长
6 months
期刊介绍: The IEEE Transactions on Knowledge and Data Engineering encompasses knowledge and data engineering aspects within computer science, artificial intelligence, electrical engineering, computer engineering, and related fields. It provides an interdisciplinary platform for disseminating new developments in knowledge and data engineering and explores the practicality of these concepts in both hardware and software. Specific areas covered include knowledge-based and expert systems, AI techniques for knowledge and data management, tools, and methodologies, distributed processing, real-time systems, architectures, data management practices, database design, query languages, security, fault tolerance, statistical databases, algorithms, performance evaluation, and applications.
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
确定
请完成安全验证×
copy
已复制链接
快去分享给好友吧!
我知道了
右上角分享
点击右上角分享
0
联系我们:info@booksci.cn Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。 Copyright © 2023 布克学术 All rights reserved.
京ICP备2023020795号-1
ghs 京公网安备 11010802042870号
Book学术文献互助
Book学术文献互助群
群 号:604180095
Book学术官方微信
小红书