[1]LIU Zhi,LI Guangdeng,LI Shiyu,et al.Modality-shared gated two-stream Transformer for visibleinfrared person re-identification[J].CAAI Transactions on Intelligent Systems,2026,21(4):979-987.[doi:10.11992/tis.202511021]
Copy
CAAI Transactions on Intelligent Systems[ISSN 1673-4785/CN 23-1538/TP] Volume:
21
Number of periods:
2026 4
Page number:
979-987
Column:
学术论文—机器感知与模式识别
Public date:
2026-07-05
- Title:
-
Modality-shared gated two-stream Transformer for visibleinfrared person re-identification
- Author(s):
-
LIU Zhi1; LI Guangdeng1; LI Shiyu1; WANG Wei2; ZHANG Xiaochuan1; 3
-
1. Liangjiang School of Artificial Intelligence, Chongqing University of Technology, Chongqing 401135, China;
2. Strategic Assessment and Consulting Center, Academy of Military Sciences, Beijing 100091, China;
3. School of Software, Chongqing Institute of Engineering, Chongqing 400056, China
-
- Keywords:
-
visible-infrared person re-identification; multi-modal learning; vision Transformer; contrastive language–image pre-training; modality-shared features extraction; modality-shared gated
- CLC:
-
TP391
- DOI:
-
10.11992/tis.202511021
- Abstract:
-
Visible-infrared person re-identification (VI-ReID) aims to match person images of the same identity across different modalities, representing a highly challenging task in the field of computer vision. In real-world scenarios, poor illumination often leads to recognition difficulties, severely limiting the practicality of person re-identification systems. Although traditional methods based on convolutional neural networks have been widely investigated, they are limited by local receptive fields and down-sampling operations, which tend to result in the loss of key modality information. To address these issues, this paper proposes a vision Transformer-based VI-ReID framework, termed modality-shared gating two-stream Transformer (MgtFormer). First, aiming at the modality discrepancy between infrared and RGB images, a grayscale image augmentation strategy is adopted to convert RGB images into grayscale ones to minimize the modality gap. Second, a modality-specific embedding module is designed to guide the ViT to extract modality-shared features more effectively. Subsequently, a modality-shared module is utilized to fully mine modality-invariant features from RGB, grayscale, and infrared images. Finally, a modality-shared gating layer is introduced to finely screen the features, retaining highly discriminative ones for subsequent re-identification tasks. MgtFormer is capable of comprehensively extracting multi-modal information and effectively preserving modality-invariant features. Extensive experiments on the SYSU-MM01 and RegDB datasets demonstrate that the proposed method outperforms current state-of-the-art techniques.