[1]王纯一,王雪.一种多模态协同学习的变电设备缺陷识别方法[J].智能系统学报,2026,21(4):1004-1012.[doi:10.11992/tis.202509009]
WANG Chunyi,WANG Xue.Multimodal collaborative learning method for defect recognition in power substation equipment[J].CAAI Transactions on Intelligent Systems,2026,21(4):1004-1012.[doi:10.11992/tis.202509009]
点击复制
《智能系统学报》[ISSN 1673-4785/CN 23-1538/TP] 卷:
21
期数:
2026年第4期
页码:
1004-1012
栏目:
学术论文—智能系统
出版日期:
2026-07-05
- Title:
-
Multimodal collaborative learning method for defect recognition in power substation equipment
- 作者:
-
王纯一1, 王雪2
-
1. 国网上海浦东供电公司电力调度控制中心, 上海 200122;
2. 西北工业大学 计算机学院, 陕西 西安 710072
- Author(s):
-
WANG Chunyi1, WANG Xue2
-
1. STATE GRID Shanghai Pudong Electric Power Supply Company, Shanghai 200122, China;
2. School of Computer Science, Northwestern Polytechnical University, Xi’an 710072, China
-
- 关键词:
-
红外图像; 紫外图像; 大语言模型; 多模态建模; 协同学习; 注意力机制; 变电设备; 缺陷识别
- Keywords:
-
infrared image; ultraviolet image; large language model; multimodal modeling; collaborative learning; attention mechanism; power substation equipment; defect recognition
- 分类号:
-
TP391.4
- DOI:
-
10.11992/tis.202509009
- 摘要:
-
为解决变电设备缺陷识别任务中多谱段图像融合特征学习不鲁棒难题,提出了一种融合语义增强与模态特征解耦的多模态协同学习方法。该方法引入具备语义理解能力的大模型DeepSeek,对红外与紫外图像的语义标注进行分析与补充,丰富标注中蕴含的语义信息表达;设计了模态共享与模态特有两个特征提取模块,通过共享查询向量和特有查询向量分别建模两种模态间的共性特征和差异特征,以提升多模态特征表示能力;引入特征对齐损失和模态解耦损失,分别用于增强模态间共享特征的一致性和抑制冗余信息干扰。实验结果表明,该方法在变电设备缺陷识别任务中优于现有主流方法,消融实验及可视化结果进一步验证了其在多模态协同学习方面的有效性。结论可为多模态融合表征学习方法及应用提供参考。
- Abstract:
-
To address the poor robustness of fused feature learning for multispectral images in substation equipment defect recognition, this study proposes a multimodal feature learning method that integrates semantic enhancement and modal feature decoupling. This method first introduces DeepSeek, a large-scale model with semantic understanding capabilities, to analyze and supplement the semantic annotations of infrared and ultraviolet images, thereby enriching the semantic information in the annotations. Modality-shared and modality-specific feature extraction modules are then designed. These modules employ shared and modality-specific query vectors to model common and distinct features across the two modalities, enhancing multimodal feature representation. Furthermore, feature alignment and decoupling losses are introduced to strengthen trans-modal shared feature consistency and suppress interference from redundant information. Extensive experimental results demonstrate that the proposed method outperforms existing mainstream approaches in electrical equipment defect recognition. Ablation studies and visualization results further validate its effectiveness in multimodal collaborative learning. This conclusion provides a reference for research and application of multimodal fusion representation learning methods.
更新日期/Last Update:
1900-01-01