Bingqian Wu , Pei An , Siwen Quan , Qiao Wu , Linjie Li , Chu’ai Zhang , Jiaqi Yang
{"title":"Rethinking the refinement stage of 3D object detection: A multi-task learning perspective with Mixture-of-Experts","authors":"Bingqian Wu , Pei An , Siwen Quan , Qiao Wu , Linjie Li , Chu’ai Zhang , Jiaqi Yang","doi":"10.1016/j.jvcir.2026.104841","DOIUrl":"10.1016/j.jvcir.2026.104841","url":null,"abstract":"<div><div>Two-stage LiDAR-based 3D object detectors have achieved state-of-the-art accuracy, yet their performance is often limited by the refinement stage. In this work, we revisit 3D object refinement from a multi-task learning perspective and identify two independent sources of negative transfer: an <strong>inter-attribute conflict</strong>, where heterogeneous regression objectives, such as center, size, and orientation, interfere during joint optimization, and an <strong>inter-sample conflict</strong>, where proposals with varying point densities lead to gradient imbalance. To address these issues, we introduce two specialized Mixture-of-Experts architectures. The <strong>Attribute-MoE</strong> decouples regression objectives into dedicated expert branches to alleviate feature conflicts, while the <strong>Sparsity-MoE</strong> employs density-aware experts to adaptively refine proposals according to point sparsity. Integrated into strong two-stage baselines, our modules consistently improve performance on the KITTI dataset and the Waymo Open Dataset. Beyond empirical gains, our analysis reveals that Attribute-MoE and Sparsity-MoE solve largely independent problems, offering a practical “toolbox” for mitigating negative transfer in 3D object refinement and advancing adaptive, task-aware detector design. Code will be released at <span><span>https://github.com/12e21/RefineMoE</span><svg><path></path></svg></span>.</div></div>","PeriodicalId":54755,"journal":{"name":"Journal of Visual Communication and Image Representation","volume":"118 ","pages":"Article 104841"},"PeriodicalIF":3.1,"publicationDate":"2026-06-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148178579","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Anchor node guided global–local graph neural networks for multimedia recommendation","authors":"Qi Ren, Desheng Cai","doi":"10.1016/j.jvcir.2026.104835","DOIUrl":"10.1016/j.jvcir.2026.104835","url":null,"abstract":"<div><div>Multimedia recommendation systems have garnered significant attention in recent years, particularly with GNN-based recommendation systems, which have demonstrated strong performance by aggregating information layer-by-layer through convolutional operations. However, previous studies have primarily focused on addressing the challenges of sparse interaction data, with limited attention given to issues arising from dense interactions. We argue that dense interactions hinder the accurate capture of user preferences, leading to bottleneck problem and an increased likelihood of noise. The introduction of GNNs further exacerbates these problems. Therefore, the problems arising from dense interaction cannot be ignored. In this paper, we propose a novel Anchor Node Guided Global–Local Graph Neural Networks for Multimedia Recommendation (AGG-LRec). To address the bottleneck problem caused by dense interactions, we propose a new graph neural network called anchor node module by using anchor nodes to aggregate the global information and interact with the target nodes. We also introduce two anchor-centric auxiliary objectives that explicitly regulate the semantic roles of anchor nodes. By enforcing anchor nodes to act as global semantic prototypes and maintaining cross-modal consistency at the anchor level, these objectives directly strengthen anchor-guided global aggregation under dense and noisy interactions. In addition, we propose a degree-sensitive edge pruning method for removing noise caused by dense interactions. This approach enhances the model’s robustness to noise by pruning edges. We also performed graph enhancement by introducing item–item edges. Experiments conducted on three datasets demonstrate that our proposed model outperforms baseline approaches. The source code is available at <span><span>https://github.com/ren3570/AGG-LRec</span><svg><path></path></svg></span>.</div></div>","PeriodicalId":54755,"journal":{"name":"Journal of Visual Communication and Image Representation","volume":"118 ","pages":"Article 104835"},"PeriodicalIF":3.1,"publicationDate":"2026-06-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148178583","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Jiaze Li , Yang Li , Xueting Ren , Lijing Zhang , Yan Qiang , Juanjuan Zhao , Huajie Yue , Yan Wang
{"title":"Spatial position association-based 3D CT reconstruction from biplanar X-rays","authors":"Jiaze Li , Yang Li , Xueting Ren , Lijing Zhang , Yan Qiang , Juanjuan Zhao , Huajie Yue , Yan Wang","doi":"10.1016/j.jvcir.2026.104845","DOIUrl":"10.1016/j.jvcir.2026.104845","url":null,"abstract":"<div><div>Computed tomography is an advanced medical imaging technique that plays a crucial role in visualizing internal structures. However, it still faces challenges such as high radiation doses and expensive equipment costs. To reduce radiation exposure and improve imaging efficiency, this paper proposes a biplanar reconstruction model called SPAX2CT, which reconstructs high-quality 3D CT volumes from 2D X-ray images. The model incorporates a global frequency-domain fusion module and a local semantic enhancement module. These components strengthen global feature boundaries and compensate for the loss of semantic information in local features. Additionally, SPAX2CT employs spatial position association encoding to capture spatial relationships within the 3D structure, addressing the insufficient spatial position correlation found in traditional methods. Experimental results show that SPAX2CT significantly improves the modeling ability for complex spatial relationships and demonstrates excellent performance on the Lung Image Database Consortium. Our model outperforms baseline methods, while also showing strong scalability and applicability.</div></div>","PeriodicalId":54755,"journal":{"name":"Journal of Visual Communication and Image Representation","volume":"118 ","pages":"Article 104845"},"PeriodicalIF":3.1,"publicationDate":"2026-06-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148178586","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Frank Sanabria-Macias , Andrés Prados-Torreblanca , Marta Marron-Romera , Javier Macias-Guarasa , José M. Buenaposada , Luis Baumela , Sira E. Palazuelos-Cagigas
{"title":"Landmark-assisted face and mouth 3D localization in smart spaces with a monocular camera","authors":"Frank Sanabria-Macias , Andrés Prados-Torreblanca , Marta Marron-Romera , Javier Macias-Guarasa , José M. Buenaposada , Luis Baumela , Sira E. Palazuelos-Cagigas","doi":"10.1016/j.jvcir.2026.104842","DOIUrl":"10.1016/j.jvcir.2026.104842","url":null,"abstract":"<div><div>Accurate <span>3D</span> mouth localization from a single camera remains challenging in smart spaces, where speakers appear at low resolution and under unconstrained head poses. Most existing audiovisual tracking methods approximate the mouth position using a fixed location within the detected face bounding box, implicitly assuming frontal pose and leading to large errors under head rotations. In addition, monocular <span>3D</span> estimation suffers from scale ambiguity, i.e., the inherent uncertainty between object size and distance in <span>2D</span> image projections. We investigate whether head pose estimation (HPE) can overcome these limitations. Two monocular approaches are proposed: <span>LPoseVL</span>, combining facial landmarks with geometric pose recovery using a rigid <span>3D</span> head model, and <span>LFreeVL</span>, based on landmark-free 6DoF pose estimation with calibration. Experiments on the <span>AV16.3</span> and <span>CAV3D</span> datasets show that explicitly modeling head pose significantly improves accuracy, achieving up to 46.9% error reduction over bounding-box-based state-of-the-art methods.</div></div>","PeriodicalId":54755,"journal":{"name":"Journal of Visual Communication and Image Representation","volume":"118 ","pages":"Article 104842"},"PeriodicalIF":3.1,"publicationDate":"2026-06-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148178585","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"IDFreq: Identity-preserved human video generation via frequency-based decomposition","authors":"Zhang Wan, Sheng Tang, Juan Cao, Yu Li","doi":"10.1016/j.jvcir.2026.104828","DOIUrl":"10.1016/j.jvcir.2026.104828","url":null,"abstract":"<div><div>Current diffusion models struggle with identity preservation in human video generation. This paper introduces IDFreq, a framework that ensures identity consistency by employing control signals in the frequency domain. We decompose facial features into low-frequency global characteristics (e.g., facial contours) and high-frequency details (e.g., unique identity markers). Specifically, a Global Facial Extractor captures the low-frequency structural information, while a Local Facial Extractor integrates fine-grained high-frequency features into the model’s Cross-Attention layers. Furthermore, we propose an optimal control-based optimization during inference to refine facial fidelity by constraining the denoising trajectory. Extensive experiments validate that IDFreq significantly improves identity consistency without compromising video quality.</div></div>","PeriodicalId":54755,"journal":{"name":"Journal of Visual Communication and Image Representation","volume":"118 ","pages":"Article 104828"},"PeriodicalIF":3.1,"publicationDate":"2026-06-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148178584","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"LWU-YOLO: A lightweight algorithm for small object detection in UAV applications","authors":"Yapeng Li , Ting Wang , Tao Li , Xin Yang","doi":"10.1016/j.jvcir.2026.104791","DOIUrl":"10.1016/j.jvcir.2026.104791","url":null,"abstract":"<div><div>Since detecting small objects in UAV imagery is challenging due to complex backgrounds and limited pixels, this paper proposes a new lightweight model based on YOLOv8s called LWU-YOLO. Initially, a task-oriented head restructuring strategy is introduced to enhance detailed feature representation, while reducing model parameters. Subsequently, an efficient multi-scale downsampling feature fusion (MDFF) module is designed to minimize the information loss during the upsampling process. Moreover, a mixed local channel attention (MLCA) mechanism is integrated into the C2f module to improve focus on critical features. Additionally, a novel Inner-PIoUv2 loss function is devised for faster convergence and higher accuracy in small object regression. Finally, experiments on the VisDrone2019 dataset show that the LWU-YOLO increases mAP@50 and mAP@50:95 by 7.3% and 4.7%, respectively, while using 55.3% fewer parameters than YOLOv8s, demonstrating an excellent balance of performance and efficiency for UAV applications.</div></div>","PeriodicalId":54755,"journal":{"name":"Journal of Visual Communication and Image Representation","volume":"117 ","pages":"Article 104791"},"PeriodicalIF":3.1,"publicationDate":"2026-04-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147600945","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Hao Qin , Zhenxue Chen , Qingqiang Guo , Q.M. Jonathan Wu , Mengxu Lu
{"title":"GGCN: Gait Recognition with Generate Network and Convolutional Neural Network","authors":"Hao Qin , Zhenxue Chen , Qingqiang Guo , Q.M. Jonathan Wu , Mengxu Lu","doi":"10.1016/j.jvcir.2026.104790","DOIUrl":"10.1016/j.jvcir.2026.104790","url":null,"abstract":"<div><div>Gait recognition is a biometric technology with wide application prospects, but it is easily affected by various covariates, which requires the gait recognition model is robust. In this paper, we design a robust gait recognition model named GGCN (Gait recognition with Generate network and Convolutional neural Network), which uses multi-type gait sequences as input and eliminates the effects of various covariates through a supervised mapping module. The GGCN processes the gait sequence in three steps. First, the generate network is used to extract low-level features and remove the features generated by interference. Then, the low-level features are input into the encoder network to obtain high-level features. Finally, the high-level features are input into the feature mapping network to acquire more recognizable features. The experimental results on the CASIA-B, OULP, and OUMVLP datasets demonstrate that our model outperforms current state-of-the-art methods.</div></div>","PeriodicalId":54755,"journal":{"name":"Journal of Visual Communication and Image Representation","volume":"117 ","pages":"Article 104790"},"PeriodicalIF":3.1,"publicationDate":"2026-04-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147600949","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Beyond bounding boxes: Segmentation supervision for robust object detection in fisheye images","authors":"Arda Oztuner, Mehmet Kilicarslan","doi":"10.1016/j.jvcir.2026.104798","DOIUrl":"10.1016/j.jvcir.2026.104798","url":null,"abstract":"<div><div>Fisheye cameras pose significant object detection challenges due to severe radial distortion, rendering traditional axis-aligned bounding boxes suboptimal for warped object shapes. We propose a pipeline that transforms bounding box annotations into instance segmentation masks using the Segment Anything Model (SAM) and validate mask fidelity against expert ground truth in both rectilinear and distorted domains. We benchmark various models on the Fisheye8K dataset, demonstrating the architectural generalizability of our approach across YOLOv8, YOLOv11, and YOLOv12. Results show that segmentation-based supervision yields substantial performance gains, improving the mean average precision (mAP@[0.5:0.95]) by up to 10 absolute points over models trained with bounding boxes, and up to 12 points in distorted outer regions. Furthermore, our approach outperforms state-of-the-art methods and establishes a new benchmark for fisheye object detection. This work highlights the specific theoretical and empirical benefits of automated segmentation-based annotation within complex, distorted imaging domains.</div></div>","PeriodicalId":54755,"journal":{"name":"Journal of Visual Communication and Image Representation","volume":"117 ","pages":"Article 104798"},"PeriodicalIF":3.1,"publicationDate":"2026-04-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147656944","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Multi-view recursive gated convolutions for 3D object recognition and retrieval","authors":"Jiangzhong Cao, Yue Cai, Huan Zhang","doi":"10.1016/j.jvcir.2026.104792","DOIUrl":"10.1016/j.jvcir.2026.104792","url":null,"abstract":"<div><div>Multi-view-based 3D shape recognition methods perform 3D object recognition and retrieval by processing series of images from various angles to generate a compact 3D descriptor. However, existing approaches often focus on integrating information from multiple views without addressing spatial interactions and redundancy when similar views are used. To overcome these challenges, we propose a novel framework, Multi-view Recursive Gated Convolutions (MVRGC). Our method first extracts features from multiple views at different scales, allowing for initial interaction of information across these views. Recursive gated convolutions are then applied to capture deeper spatial reciprocity and fine-tune feature interactions among views. Additionally, a preferred view module is introduced to reduce view redundancy by favoring distinctive and representative views. This module selects a subset of views that best describe the object while minimizing overlap. Experimental results on shape benchmark datasets demonstrate that MVRGC outperforms existing methods in 3D object recognition and retrieval tasks.</div></div>","PeriodicalId":54755,"journal":{"name":"Journal of Visual Communication and Image Representation","volume":"117 ","pages":"Article 104792"},"PeriodicalIF":3.1,"publicationDate":"2026-04-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147600946","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Densely aggregated U-net with spatial-spectral interaction transformer for hyperspectral compressed imaging reconstruction","authors":"Yun-Hui Li","doi":"10.1016/j.jvcir.2026.104795","DOIUrl":"10.1016/j.jvcir.2026.104795","url":null,"abstract":"<div><div>Hyperspectral imaging offers critical spectral information for applications such as material analysis and camouflage recognition. However, the acquisition of hyperspectral data cubes is inherently constrained by the Nyquist sampling theorem. While compressed sensing theory enables snapshot imaging by compressing the data cube into a 2D measurement, the ill-posed reconstruction remains a significant challenge. Recent deep learning methods, particularly vision transformers, have advanced the state-of-the-art (SOTA). Despite this, existing networks typically employ spectral or spatial self-attentions in isolation, blindly pursuing a global receptive field at the cost of computational efficiency and representational flexibility. Additionally, the vanilla skip connection in U-Nets is insufficient for effective multi-scale information transmission between encoder and decoder. To address these issues, we propose a Densely aggregated U-Net with a Spatial-Spectral Interaction Transformer (DSST). DSST parallelizes patch-based spectral self-attention and window-based spatial self-attention, complemented by an interaction mechanism. Furthermore, it introduces a densely aggregated skip connection to collect multi-scale features and bridge the semantic gap. Experimental results on both simulated and real-world scenes demonstrate that DSST achieves competitive performance with lower computational and memory costs compared to other end-to-end networks. Moreover, it offers faster inference speeds than deep unfolding networks.</div></div>","PeriodicalId":54755,"journal":{"name":"Journal of Visual Communication and Image Representation","volume":"117 ","pages":"Article 104795"},"PeriodicalIF":3.1,"publicationDate":"2026-04-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147600948","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}