IEEE Transactions on Circuits and Systems for Video Technology最新文献

筛选
英文 中文
Deep-JGAC: End-to-End Deep Joint Geometry and Attribute Compression for Dense Colored Point Clouds Deep- jgac:密集彩色点云的端到端深度节点几何和属性压缩
IF 10.8 1区 工程技术
IEEE Transactions on Circuits and Systems for Video Technology Pub Date : 2026-08-01 Epub Date: 2026-03-20 DOI: 10.1109/TCSVT.2026.3676115
Yun Zhang;Zhiwei Guo;Zixi Guo;Linwei Zhu;C.-C. Jay Kuo
{"title":"Deep-JGAC: End-to-End Deep Joint Geometry and Attribute Compression for Dense Colored Point Clouds","authors":"Yun Zhang;Zhiwei Guo;Zixi Guo;Linwei Zhu;C.-C. Jay Kuo","doi":"10.1109/TCSVT.2026.3676115","DOIUrl":"https://doi.org/10.1109/TCSVT.2026.3676115","url":null,"abstract":"Colored point cloud becomes a fundamental representation in the realm of 3D vision. Effective Point Cloud Compression (PCC) is urgently needed due to the huge amount of data. In this paper, we propose an end-to-end Deep Joint Geometry and Attribute Compression (Deep-JGAC) method for dense colored point clouds. First, we propose a flexible Deep-JGAC framework, where the geometry and attribute encoders are compatible with either learning or non-learning encoders. Second, we propose an end-to-end deep residual self-attention-based geometry encoder to improve geometry coding efficiency, where a Hybrid Residual Self-attention Module (HRSM) is proposed to enhance geometry representation by considering its geometrical importance. Third, to solve the mismatch between the point cloud geometry and attribute caused by the geometry compression distortion, we propose an optimized re-colorization module to attach attribute to the geometrically distorted point cloud for attribute coding, which lowers the computational complexity. Extensive experimental results demonstrate that, in terms of the geometry quality metric D1-PSNR, the proposed Deep-JGAC achieves average Bjøntegaard Delta Bit Rate (BDBR) of −82.96%, −44.63%, −36.46%, -41.72%, and -31.16% compared to the G-PCC (Octree), G-PCC (Trisoup), V-PCC, GRASP, and PCGCv2, respectively. For the perceptual joint quality metric MS-GraphSIM, Deep-JGAC achieves an average BDBR of −48.72%, −57.14%, −14.67% and -13.37% against G-PCC(Octree), IT-DL-PCC, V-PCC, and DeepPCC, respectively. In addition, the costs of encoding/decoding time are reduced by 32.8%/30.8%, 80.1%/81.8%, 97.2%/35.7%, 98.4%/92.3%, and 96.4%/99.6% on average compared to G-PCC (Octree), G-PCC (Trisoup), V-PCC, IT-DL-PCC and DeepPCC. The code and pre-trained models are available at <uri>https://github.com/SYSU-Video/Deep-JGAC</uri>","PeriodicalId":13082,"journal":{"name":"IEEE Transactions on Circuits and Systems for Video Technology","volume":"36 8","pages":"11416-11430"},"PeriodicalIF":10.8,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148675597","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":1,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
3 × 3 Kernel Is All You Need for Vision 3 × 3内核是所有你需要的视觉
IF 10.8 1区 工程技术
IEEE Transactions on Circuits and Systems for Video Technology Pub Date : 2026-08-01 Epub Date: 2026-04-21 DOI: 10.1109/TCSVT.2026.3686189
Shenqi Lai;Mengjian Li;Haifeng Liu;Xueming Qian;Deng Cai;Yaxiong Wang
{"title":"3 × 3 Kernel Is All You Need for Vision","authors":"Shenqi Lai;Mengjian Li;Haifeng Liu;Xueming Qian;Deng Cai;Yaxiong Wang","doi":"10.1109/TCSVT.2026.3686189","DOIUrl":"https://doi.org/10.1109/TCSVT.2026.3686189","url":null,"abstract":"Most modern Convolutional Neural Networks (CNNs) employ a multi-branch structure with various-sized convolutions to capture long- and short-range dependencies. However, these CNNs use large kernel convolutions (e.g., astonishingly 101 kernels) and specialized techniques (e.g., re-parameterization and sparsity), increasing complexity in both training and inference stages. This paper focuses on designing an efficient CNN based on pure <inline-formula> <tex-math>$3times 3$ </tex-math></inline-formula> convolutions without introducing complex operations and techniques. Specifically, we propose a Spatial Pyramid (SP) block, which consists of the Multi-branch Residual (MbR) module and the Gated-branch Residual (GbR) module. The MbR introduces multiscale pooling as the key component, thus capturing long-range visual cues through large down-sampling rates and shorter-range dependencies through low down-sampling rates while maintaining low computational complexity. Besides, the GbR uses one <inline-formula> <tex-math>$3times 3$ </tex-math></inline-formula> convolution to refine dependencies along spatial and channel dimensions. Based on the SP block, we construct the Spatial Pyramid CNN (SPCNN), a model composed exclusively of Point-Wise Convolution and <inline-formula> <tex-math>$3times 3$ </tex-math></inline-formula> Depth-Wise Convolution. Under comparable computational complexity, SPCNN significantly outperforms the state-of-the-art CNN PeLK (83.6% vs 82.6%) with only <inline-formula> <tex-math>$3times 3$ </tex-math></inline-formula> kernels (compared to <inline-formula> <tex-math>$101times 101$ </tex-math></inline-formula> kernels in PeLK). Besides, our SPCNN demonstrates comparability with state-of-the-art backbones in lightweight models, object detection, instance segmentation, and semantic segmentation. Moreover, evaluations of four image retrieval benchmarks also demonstrate the effectiveness. All codes are released at <uri>https://github.com/xiaolai-sqlai/SPCNN</uri>","PeriodicalId":13082,"journal":{"name":"IEEE Transactions on Circuits and Systems for Video Technology","volume":"36 8","pages":"11363-11375"},"PeriodicalIF":10.8,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148675598","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":1,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
Double Buffer Vaccination: Bolstering Immunity Against Catastrophic Forgetting in Continual Medical Image Segmentation 双重缓冲免疫:增强连续医学图像分割中对灾难性遗忘的免疫力
IF 10.8 1区 工程技术
IEEE Transactions on Circuits and Systems for Video Technology Pub Date : 2026-08-01 Epub Date: 2026-04-09 DOI: 10.1109/TCSVT.2026.3682343
Kai Chen;Hewei Wang;Pinzhuo Tian;Jing Huo;Taihang Zhen;Yang Gao
{"title":"Double Buffer Vaccination: Bolstering Immunity Against Catastrophic Forgetting in Continual Medical Image Segmentation","authors":"Kai Chen;Hewei Wang;Pinzhuo Tian;Jing Huo;Taihang Zhen;Yang Gao","doi":"10.1109/TCSVT.2026.3682343","DOIUrl":"https://doi.org/10.1109/TCSVT.2026.3682343","url":null,"abstract":"Deep learning in medical imaging is crucial for clinical diagnosis and treatment guidance, gaining attention due to its ability to capture essential disease characteristics. However, continuous advancements in imaging technology, scanner variety, and varying skills among practitioners often hinder the practical application of machine learning. These factors lead to models becoming obsolete due to shifts in the imaging domain, resulting in decreased predictive performance, a phenomenon known as “catastrophic forgetting” in continual segmentation. We propose a continual learning method to address the degradation in model segmentation performance caused by these cross-domain shifts. Our approach utilizes a double buffer strategy comprising a typical replay buffer and its augmented version, the virtual replay buffer. The virtual replay buffer is derived from the typical one through Continuous Frequency Space Interpolation, collaboratively storing a minimal number of previously learned images. This strategy allows the model to revisit a small number of past images while learning in a new domain, thereby proactively bridging distribution gaps. Additionally, we enforce consistency constraints between the typical replay buffer and its augmented virtual data, effectively reducing the structural feature distance across different imaging domains. Finally, we further constrain the model using a loose deep knowledge distillation strategy to explicitly guide gradient updates. We tested our method on continuously obtained prostate MRI data from six research institutes, retinal fundus image data from four different clinical centers, and the CHAOS dataset. The results demonstrate the consistent superiority of our approach. Our code is publicly available at: <uri>https://github.com/KaiChenNJ/DBV</uri>","PeriodicalId":13082,"journal":{"name":"IEEE Transactions on Circuits and Systems for Video Technology","volume":"36 8","pages":"12545-12559"},"PeriodicalIF":10.8,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148675693","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":1,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
MOFM: A Multiple-in-One Flow Mamba for Unregistered Multi-Modal Image Fusion MOFM:一个多合一流曼巴未注册的多模态图像融合
IF 10.8 1区 工程技术
IEEE Transactions on Circuits and Systems for Video Technology Pub Date : 2026-08-01 Epub Date: 2026-04-07 DOI: 10.1109/TCSVT.2026.3681540
Bo Yang;Zhaohui Jiang;Dong Pan;Zhiping Lin;Weihua Gui
{"title":"MOFM: A Multiple-in-One Flow Mamba for Unregistered Multi-Modal Image Fusion","authors":"Bo Yang;Zhaohui Jiang;Dong Pan;Zhiping Lin;Weihua Gui","doi":"10.1109/TCSVT.2026.3681540","DOIUrl":"https://doi.org/10.1109/TCSVT.2026.3681540","url":null,"abstract":"Multi-modal image fusion aims to integrate complementary cues from different modalities into a single image, facilitating downstream tasks such as object detection. However, input image pairs are often misaligned due to rigid or non-rigid deformations in practical scenarios. Such misregistration introduces structural distortions and visual artifacts, reducing the reliability of the fused results and limiting their effectiveness for downstream vision applications. While existing methods demonstrate satisfactory results under specific deformation scenarios, they exhibit limited generalization to diverse and severe misregistrations. To this end, this study proposes a unified multiple-in-one flow Mamba framework for registering various image deformations and generating high-quality fused results. Specifically, a hierarchical flow Mamba is designed to model rigid and non-rigid flow fields and enhance adaptability to complex deformations by progressively refining misaligned features. To better distinguish between rigid and non-rigid types, a flow field classifier predicts rigid/non-rigid categories and provides prompts for high-level feature modulation. Furthermore, a flow-guided fusion Mamba module is developed to aggregate aligned multi-level modality features and generate fused images, while an iterative training strategy enables collaborative optimization by using fusion outputs to refine flow estimation. Experiments across three representative modality tasks demonstrate that the proposed method delivers superior fusion performance while maintaining applicability to object detection. The code will be available at: <uri>https://github.com/BOYang-pro/MOFM</uri>","PeriodicalId":13082,"journal":{"name":"IEEE Transactions on Circuits and Systems for Video Technology","volume":"36 8","pages":"12296-12310"},"PeriodicalIF":10.8,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148675730","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":1,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
IEEE Circuits and Systems Society Information IEEE电路与系统学会信息
IF 10.8 1区 工程技术
IEEE Transactions on Circuits and Systems for Video Technology Pub Date : 2026-08-01 Epub Date: 2026-08-05 DOI: 10.1109/TCSVT.2026.3716126
{"title":"IEEE Circuits and Systems Society Information","authors":"","doi":"10.1109/TCSVT.2026.3716126","DOIUrl":"https://doi.org/10.1109/TCSVT.2026.3716126","url":null,"abstract":"","PeriodicalId":13082,"journal":{"name":"IEEE Transactions on Circuits and Systems for Video Technology","volume":"36 8","pages":"C3-C3"},"PeriodicalIF":10.8,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=11643477","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148675749","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":1,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
Generating Multi-Modal Knowledge Clues as an Image: Toward Improving Image-Sequence Reasoning With Assisted Visual Input 以图像形式生成多模态知识线索:利用辅助视觉输入改进图像序列推理
IF 10.8 1区 工程技术
IEEE Transactions on Circuits and Systems for Video Technology Pub Date : 2026-08-01 Epub Date: 2026-03-20 DOI: 10.1109/TCSVT.2026.3676204
Guanghui Ye;Huan Zhao;Yixian Shen;Jiaqi Li;Fengnan Li;Zhihua Jiang;Keqin Li
{"title":"Generating Multi-Modal Knowledge Clues as an Image: Toward Improving Image-Sequence Reasoning With Assisted Visual Input","authors":"Guanghui Ye;Huan Zhao;Yixian Shen;Jiaqi Li;Fengnan Li;Zhihua Jiang;Keqin Li","doi":"10.1109/TCSVT.2026.3676204","DOIUrl":"https://doi.org/10.1109/TCSVT.2026.3676204","url":null,"abstract":"Recent multi-modal large language models (MLLMs) have exhibited powerful abilities in addressing complex vision-language tasks such as image-sequence reasoning (ISR). However, significant challenges remain, e.g., it is still difficult for the MLLMs to fully capture and represent cross-image visual knowledge such as scene relations, attributes, and entity links between multiple images, which hinders them from better solving ISR. To alleviate these issues, we introduce a novel concept <bold>Vi</b>suali<bold>z</b>ed <bold>K</b>nowledge <bold>C</b>lue (<bold>VizKC</b>) - synthetic images that encode key visual and external knowledge from a sequence of input images and are then used alongside the original input images within a multi-image MLLM to enhance reasoning performance. Accordingly, we propose an accompanying approach named <bold>VizKC-ISR</b>, composed of two modules - VizKC generation and VizKC utilization. Specifically, in the generation module, VizKC-ISR follows a <italic>See-Find-Fuse</i> pipeline: 1) “<italic>See - Scene Perception</i>”, to construct an initial VizKC that incorporates scene relations of key visual entities detected from an original image; 2) “<italic>Find - Knowledge Generation</i>”, to generate enriched image captions with real-world knowledge and fine-grained entity details and then extract structured knowledge tuples from generated captions; 3) “<italic>Fuse - Image Editing</i>”, to introduce relevant knowledge tuples into the VizKC via iterative image editing. In the utilization module, we employ a multi-image MLLM (e.g., mPLUG-Owl3) to solve the VizKC-assisted ISR tasks by reasoning with generated knowledge clues. We evaluate VizKC-ISR on nine ISR benchmarks categorized into three multi-image scenarios. The results show that our VizKC-ISR performs best in all tasks, e.g., obtaining the highest average accuracy of 63.1% and surpassing the mPLUG-Owl3 baseline by 6.4 absolute points, due to the bridge between visually-grounded reasoning and multi-modal knowledge challenges.","PeriodicalId":13082,"journal":{"name":"IEEE Transactions on Circuits and Systems for Video Technology","volume":"36 8","pages":"11197-11214"},"PeriodicalIF":10.8,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148675754","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":1,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and Understanding MotionGPT-2:一个用于运动生成和理解的通用运动语言模型
IF 10.8 1区 工程技术
IEEE Transactions on Circuits and Systems for Video Technology Pub Date : 2026-08-01 Epub Date: 2026-03-19 DOI: 10.1109/TCSVT.2026.3694713
Yuan Wang;Di Huang;Yaqi Zhang;Wanli Ouyang;Jile Jiao;Xuetao Feng;Dan Xu;Shixiang Tang
{"title":"MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and Understanding","authors":"Yuan Wang;Di Huang;Yaqi Zhang;Wanli Ouyang;Jile Jiao;Xuetao Feng;Dan Xu;Shixiang Tang","doi":"10.1109/TCSVT.2026.3694713","DOIUrl":"https://doi.org/10.1109/TCSVT.2026.3694713","url":null,"abstract":"Generating lifelike human motions from descriptive texts has experienced remarkable research focus in recent years, propelled by the emerging requirements of digital humans. Despite impressive advances, existing approaches are often constrained by limited control modalities, task specificity, and focus solely on body motion representations. In this paper, we present MotionGPT-2, a unified Large Motion-Language Model (LMLM) that addresses these limitations. MotionGPT-2 accommodates multiple motion-relevant tasks and supports multimodal control conditions through pre-trained Large Language Models (LLMs). It quantizes multimodal inputs—such as text and single-frame poses—into discrete, LLM-interpretable tokens, seamlessly integrating them into the LLM’s vocabulary. These tokens are then organized into unified prompts, guiding the LLM to generate motion outputs through a pretraining-then-finetuning paradigm. We also show that the proposed MotionGPT-2 is highly adaptable to the challenging 3D holistic motion generation task, enabled by the innovative motion discretization framework, Part-Aware VQVAE, which facilitates fine-grained representations of body and hand movements. Extensive experiments and visualizations validate the effectiveness of our method, demonstrating the adaptability of MotionGPT-2 across motion generation, motion captioning, and generalized motion completion tasks.","PeriodicalId":13082,"journal":{"name":"IEEE Transactions on Circuits and Systems for Video Technology","volume":"36 8","pages":"11027-11041"},"PeriodicalIF":10.8,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=11527021","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148675812","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":1,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
HPE-CogVLM: Advancing Vision Language Models With a Head Pose Grounding Task HPE-CogVLM:基于头部姿势基础任务的视觉语言模型
IF 10.8 1区 工程技术
IEEE Transactions on Circuits and Systems for Video Technology Pub Date : 2026-08-01 Epub Date: 2026-03-20 DOI: 10.1109/TCSVT.2026.3675940
Yu Tian;Tianqi Shao;Tsukasa Demizu;Xuyang Wu;Hsin-Tai Wu
{"title":"HPE-CogVLM: Advancing Vision Language Models With a Head Pose Grounding Task","authors":"Yu Tian;Tianqi Shao;Tsukasa Demizu;Xuyang Wu;Hsin-Tai Wu","doi":"10.1109/TCSVT.2026.3675940","DOIUrl":"https://doi.org/10.1109/TCSVT.2026.3675940","url":null,"abstract":"Head pose estimation (HPE) requires a sophisticated understanding of 3D spatial relationships to generate precise yaw, pitch, and roll angles. Previous HPE models, primarily CNN-based, rely on cropped close-up human head images as inputs and often lack robustness in real-world scenario. Vision Language Models (VLMs) can analyze entire images while focusing on specific objects through their attention mechanisms. In this paper, we propose a novel framework to improve the HPE accuracy by leveraging the object detection grounding capability of a VLM, referred to as CogVLM. We empirically find that directly LoRA fine-tuning of this VLM for the HPE task fails to achieve desirable HPE accuracy, while some model merging methods can improve accuracy but frequently produce blended invalid response formats, struggling to handle both object detection and HPE tasks simultaneously. To integrate HPE capability into CogVLM effectively, we develop a novel LoRA layer-based model merging method. This merging approach applies a high cosine similarity threshold and a “winner-takes-all” layer selection strategy, aligning attention to the HPE task while preserving original object detection knowledge. It successfully resolves issues with blended invalid response formats and improves accuracy. Results show that our HPE-CogVLM achieves a 31.5% reduction in Mean Absolute Error over the current state-of-the-art CNN model, 6DRepNet, in cross-dataset evaluation. Furthermore, HPE-CogVLM outperforms both directly LoRA fine-tuned and task arithmetic-based merged VLMs across all HPE metrics.","PeriodicalId":13082,"journal":{"name":"IEEE Transactions on Circuits and Systems for Video Technology","volume":"36 8","pages":"10911-10926"},"PeriodicalIF":10.8,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=11449295","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148675851","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":1,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
Interaction-Driven Edge Crisping for Underwater Salient Object Detection 水下显著目标检测的交互驱动边缘卷曲
IF 10.8 1区 工程技术
IEEE Transactions on Circuits and Systems for Video Technology Pub Date : 2026-08-01 Epub Date: 2026-03-31 DOI: 10.1109/TCSVT.2026.3679543
Zetian Mi;Shuaiyong Jiang;Yuanyuan Li;Guanxi Li;Jiqing Zhang;Huibing Wang;Xianping Fu
{"title":"Interaction-Driven Edge Crisping for Underwater Salient Object Detection","authors":"Zetian Mi;Shuaiyong Jiang;Yuanyuan Li;Guanxi Li;Jiqing Zhang;Huibing Wang;Xianping Fu","doi":"10.1109/TCSVT.2026.3679543","DOIUrl":"https://doi.org/10.1109/TCSVT.2026.3679543","url":null,"abstract":"Underwater salient object detection (USOD) faces greater challenges than general scenes due to the edge blurring which is caused by light absorption and scattering in water. Existing methods employ unrefined edge feature to perform unidirectional guidance on saliency feature, resulting in the coarse edge of saliency map. To address this issue, we propose a novel interaction-driven edge crisping network (IDENet) for underwater salient object detection. IDENet facilitates the bidirectional modulation of inter-features and the self-refinement of intra-feature, generates crisp saliency map and edge map. In IDENet, the interaction-driven edge guidance module (IDEGM) is designed to utilize cross-feature interaction by leveraging their correlations, facilitating saliency feature’s awareness of edge information, mitigating the interference of non-salient objects in edge feature. To learn more accurate edge region of the salient object, the edge intersection-and-union loss function (EIUL) is introduced to restrict the intersection and union of predicted saliency maps and edge maps to prevent over-expansion or under-contraction. Experimental results on two latest underwater datasets demonstrate the superiority of the proposed method over the state-of-the-art models. The source code of our method will be made available at <uri>https://github.com/UnderwaterVisionMZTdlmu/IDENet</uri>","PeriodicalId":13082,"journal":{"name":"IEEE Transactions on Circuits and Systems for Video Technology","volume":"36 8","pages":"11532-11546"},"PeriodicalIF":10.8,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148675900","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":1,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
Free Your Hands: Human-Demonstration-Free Diffusion Policy Adaptation for Robotic Manipulation 解放你的双手:机器人操作的无人类示范扩散策略适应
IF 10.8 1区 工程技术
IEEE Transactions on Circuits and Systems for Video Technology Pub Date : 2026-08-01 Epub Date: 2026-04-21 DOI: 10.1109/TCSVT.2026.3686290
Ge Yuan;Ming Yang;Jing Zhang;Jianxin Pang;Dong Xu
{"title":"Free Your Hands: Human-Demonstration-Free Diffusion Policy Adaptation for Robotic Manipulation","authors":"Ge Yuan;Ming Yang;Jing Zhang;Jianxin Pang;Dong Xu","doi":"10.1109/TCSVT.2026.3686290","DOIUrl":"https://doi.org/10.1109/TCSVT.2026.3686290","url":null,"abstract":"Diffusion-based policies demonstrate remarkable capabilities in generating robust and precise actions for embodied tasks. However, when adapting to a new environment with visual domain gap such as lighting variation, object appearance change, different background, and unseen distractors, these methods require a substantial amount of human demonstrations gathered through teleoperation. This data collection process incurs significant costs in terms of both time and financial resources. To this end, we introduce <italic>FreeHand</i>, which enables cross-environment adaptation without requiring additional teleoperated demonstrations, <italic>freeing human hands</i> from the time-consuming data collection process. Specifically, we find that a well-trained diffusion policy network struggles to adapt to new environments due to distribution mismatches at both the visual and policy levels. Simply aligning domains at the visual level is insufficient, as even subtle visual changes in the environment can lead to severe action failures. Therefore, we propose a two-level alignment scheme. The visual-level alignment is achieved through adversarial training between the visual features from old and new environments. For the policy-level alignment, we introduce a novel noise-aware adversarial learning strategy for the diffusion policy. We validate our approach through comprehensive cross-domain adaptation experiments under three settings: Sim2Sim (PushT, CALVIN, RoboTwin), Real2Sim (SimplerEnv), and Real2Real (real-world robotic deployment). Our method demonstrates significant improvement with minimal additional parameter overhead, showcasing its effectiveness and efficiency.","PeriodicalId":13082,"journal":{"name":"IEEE Transactions on Circuits and Systems for Video Technology","volume":"36 8","pages":"12586-12600"},"PeriodicalIF":10.8,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148675993","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":1,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
0
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
确定
请完成安全验证×
相关产品
×
本文献相关产品
联系我们:info@booksci.cn Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。 Copyright © 2023 布克学术 All rights reserved.
京ICP备2023020795号-1
ghs 京公网安备 11010802042870号
Book学术文献互助
Book学术文献互助群
群 号:604180095
Book学术官方微信
小红书