Zhiyu Guan, Zhaofa Wang, Gan Zhang, Luwei Li, Miaomiao Zhang, Zhiping Shi, Na Jiang
{"title":"Multi-object tracking review: retrospective and emerging trend","authors":"Zhiyu Guan, Zhaofa Wang, Gan Zhang, Luwei Li, Miaomiao Zhang, Zhiping Shi, Na Jiang","doi":"10.1007/s10462-025-11212-y","DOIUrl":null,"url":null,"abstract":"<div><p>Multi-object tracking (MOT) is a critical task involving detecting and continuously tracking multiple objects within a video sequence. It is widely used in various fields, such as autonomous driving and intelligent security. In recent years, deep learning architectures have effectively promoted the development of MOT. However, this task poses significant challenges regarding accuracy due to occlusion/truncation, light variation, camera movement. Researchers have proposed many methods to address these issues to reduce trajectory fragmentation, identity switches, and missing targets. To better understand these advancements, it is essential to categorize the approaches based on their methodologies. This article reviewed the recent development of MOT, divided into Tracking by Detection (TBD) and End-to-End (E2E). By introducing and comparing the two types of tracking algorithms, readers can quickly understand the current development status of MOT. Meanwhile, this review summarizes the links to open-source code of excellent algorithms and common benchmark datasets in the appendix. And provide a unified MOT toolkit that includes evaluation and visualization at https://github.com/guanzhiyu817/MOT-tools. In addition, this review discusses the future directions of MOT, specifically cross-modal reasoning.</p></div>","PeriodicalId":8449,"journal":{"name":"Artificial Intelligence Review","volume":"58 8","pages":""},"PeriodicalIF":10.7000,"publicationDate":"2025-05-07","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://link.springer.com/content/pdf/10.1007/s10462-025-11212-y.pdf","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Artificial Intelligence Review","FirstCategoryId":"94","ListUrlMain":"https://link.springer.com/article/10.1007/s10462-025-11212-y","RegionNum":2,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q1","JCRName":"COMPUTER SCIENCE, ARTIFICIAL INTELLIGENCE","Score":null,"Total":0}
引用次数: 0
Abstract
Multi-object tracking (MOT) is a critical task involving detecting and continuously tracking multiple objects within a video sequence. It is widely used in various fields, such as autonomous driving and intelligent security. In recent years, deep learning architectures have effectively promoted the development of MOT. However, this task poses significant challenges regarding accuracy due to occlusion/truncation, light variation, camera movement. Researchers have proposed many methods to address these issues to reduce trajectory fragmentation, identity switches, and missing targets. To better understand these advancements, it is essential to categorize the approaches based on their methodologies. This article reviewed the recent development of MOT, divided into Tracking by Detection (TBD) and End-to-End (E2E). By introducing and comparing the two types of tracking algorithms, readers can quickly understand the current development status of MOT. Meanwhile, this review summarizes the links to open-source code of excellent algorithms and common benchmark datasets in the appendix. And provide a unified MOT toolkit that includes evaluation and visualization at https://github.com/guanzhiyu817/MOT-tools. In addition, this review discusses the future directions of MOT, specifically cross-modal reasoning.
期刊介绍:
Artificial Intelligence Review, a fully open access journal, publishes cutting-edge research in artificial intelligence and cognitive science. It features critical evaluations of applications, techniques, and algorithms, providing a platform for both researchers and application developers. The journal includes refereed survey and tutorial articles, along with reviews and commentary on significant developments in the field.