[1]WANG Chunyi,WANG Xue.Multimodal collaborative learning method for defect recognition in power substation equipment[J].CAAI Transactions on Intelligent Systems,2026,21(4):1004-1012.[doi:10.11992/tis.202509009]
Copy
CAAI Transactions on Intelligent Systems[ISSN 1673-4785/CN 23-1538/TP] Volume:
21
Number of periods:
2026 4
Page number:
1004-1012
Column:
学术论文—智能系统
Public date:
2026-07-05
- Title:
-
Multimodal collaborative learning method for defect recognition in power substation equipment
- Author(s):
-
WANG Chunyi1; WANG Xue2
-
1. STATE GRID Shanghai Pudong Electric Power Supply Company, Shanghai 200122, China;
2. School of Computer Science, Northwestern Polytechnical University, Xi’an 710072, China
-
- Keywords:
-
infrared image; ultraviolet image; large language model; multimodal modeling; collaborative learning; attention mechanism; power substation equipment; defect recognition
- CLC:
-
TP391.4
- DOI:
-
10.11992/tis.202509009
- Abstract:
-
To address the poor robustness of fused feature learning for multispectral images in substation equipment defect recognition, this study proposes a multimodal feature learning method that integrates semantic enhancement and modal feature decoupling. This method first introduces DeepSeek, a large-scale model with semantic understanding capabilities, to analyze and supplement the semantic annotations of infrared and ultraviolet images, thereby enriching the semantic information in the annotations. Modality-shared and modality-specific feature extraction modules are then designed. These modules employ shared and modality-specific query vectors to model common and distinct features across the two modalities, enhancing multimodal feature representation. Furthermore, feature alignment and decoupling losses are introduced to strengthen trans-modal shared feature consistency and suppress interference from redundant information. Extensive experimental results demonstrate that the proposed method outperforms existing mainstream approaches in electrical equipment defect recognition. Ablation studies and visualization results further validate its effectiveness in multimodal collaborative learning. This conclusion provides a reference for research and application of multimodal fusion representation learning methods.