[1]刘智,李广登,李诗雨,等.用于可见–红外行人重识别的模态共享门控双流Transformer[J].智能系统学报,2026,21(4):979-987.[doi:10.11992/tis.202511021]
LIU Zhi,LI Guangdeng,LI Shiyu,et al.Modality-shared gated two-stream Transformer for visibleinfrared person re-identification[J].CAAI Transactions on Intelligent Systems,2026,21(4):979-987.[doi:10.11992/tis.202511021]
点击复制
《智能系统学报》[ISSN 1673-4785/CN 23-1538/TP] 卷:
21
期数:
2026年第4期
页码:
979-987
栏目:
学术论文—机器感知与模式识别
出版日期:
2026-07-05
- Title:
-
Modality-shared gated two-stream Transformer for visibleinfrared person re-identification
- 作者:
-
刘智1, 李广登1, 李诗雨1, 王伟2, 张小川1,3
-
1. 重庆理工大学 两江人工智能学院, 重庆 401135;
2. 军事科学院 战略评估咨询中心, 北京 100091;
3. 重庆工程学院 软件学院, 重庆 400056
- Author(s):
-
LIU Zhi1, LI Guangdeng1, LI Shiyu1, WANG Wei2, ZHANG Xiaochuan1,3
-
1. Liangjiang School of Artificial Intelligence, Chongqing University of Technology, Chongqing 401135, China;
2. Strategic Assessment and Consulting Center, Academy of Military Sciences, Beijing 100091, China;
3. School of Software, Chongqing Institute of Engineering, Chongqing 400056, China
-
- 关键词:
-
可见–红外行人重识别; 多模态学习; 视觉Transformer; 对比语言-图像预训练; 模态共享特征提取; 模态共享门控
- Keywords:
-
visible-infrared person re-identification; multi-modal learning; vision Transformer; contrastive language–image pre-training; modality-shared features extraction; modality-shared gated
- 分类号:
-
TP391
- DOI:
-
10.11992/tis.202511021
- 摘要:
-
可见光–红外行人重识别旨在匹配不同模态下的同一身份行人图像,是计算机视觉领域一项极具挑战性的任务。在现实场景中,光照不足往往导致识别困难,严重限制了行人重识别系统的实用性。尽管基于卷积神经网络的传统方法已得到广泛研究,但受限于局部感受野及下采样操作,这类方法容易造成模态信息的丢失。针对上述问题,本文提出了一种基于vision Transformer(ViT)的可见光–红外行人重识框架——模态共享门控双流Transformer。针对红外与RGB图像间的模态差异,采用灰度图像增强策略,将RGB图像转换为灰度图以最小化模态间隙。设计模态特有嵌入模块,引导ViT更有效地提取模态共享特征。随后,利用模态共享模块充分挖掘RGB、灰度及红外3类图像中的模态无关特征。最后,引入模态共享门控层对特征进行精细筛选,保留高鉴别力特征以用于后续重识别任务。本文框架能够全面提取多模态信息,并有效保留模态无关特征。在SYSU-MM01和RegDB数据集上的大量实验表明,该方法优于当前主流技术。
- Abstract:
-
Visible-infrared person re-identification (VI-ReID) aims to match person images of the same identity across different modalities, representing a highly challenging task in the field of computer vision. In real-world scenarios, poor illumination often leads to recognition difficulties, severely limiting the practicality of person re-identification systems. Although traditional methods based on convolutional neural networks have been widely investigated, they are limited by local receptive fields and down-sampling operations, which tend to result in the loss of key modality information. To address these issues, this paper proposes a vision Transformer-based VI-ReID framework, termed modality-shared gating two-stream Transformer (MgtFormer). First, aiming at the modality discrepancy between infrared and RGB images, a grayscale image augmentation strategy is adopted to convert RGB images into grayscale ones to minimize the modality gap. Second, a modality-specific embedding module is designed to guide the ViT to extract modality-shared features more effectively. Subsequently, a modality-shared module is utilized to fully mine modality-invariant features from RGB, grayscale, and infrared images. Finally, a modality-shared gating layer is introduced to finely screen the features, retaining highly discriminative ones for subsequent re-identification tasks. MgtFormer is capable of comprehensively extracting multi-modal information and effectively preserving modality-invariant features. Extensive experiments on the SYSU-MM01 and RegDB datasets demonstrate that the proposed method outperforms current state-of-the-art techniques.
备注/Memo
收稿日期:2025-11-15。
作者简介:刘智,副教授,主要研究方向为计算机视觉、机器学习、视频与信号分析。主持或参与国家自然科学基金、重庆市自然科学基金等纵向及企业横向项目20余项。获国家发明专利授权4项,以第一或通讯作者发表学术论文20余篇。E-mail:liuzhi@cqut.edu.cn。;李广登,硕士,主要研究方向为计算机视觉和行人重识别。E-mail:guangdeng.lee@gmail.com。;张小川,教授,CAAI杰出会员,CAAI机器博弈专委会主任委员,重庆工程学院智能系统工程中心主任。主要研究方向为软件工程、机器博弈、自然语言处理、机器学习和智能机器人。主持和参与纵向项目38项,获省部级自然科学奖等2项,出版专著和 教材5部,发表论文100余篇。E-mail: cqpczxc@qq.com。
通讯作者:张小川. E-mail:cqpczxc@qq.com
更新日期/Last Update:
1900-01-01