[1]SHEN Xueli,YU Jiahao.Multimodal graph contrastive learning recommendation based on fusion diffusion models[J].CAAI Transactions on Intelligent Systems,2026,21(4):1044-1054.[doi:10.11992/tis.202507015]
Copy
CAAI Transactions on Intelligent Systems[ISSN 1673-4785/CN 23-1538/TP] Volume:
21
Number of periods:
2026 4
Page number:
1044-1054
Column:
人工智能院长论坛
Public date:
2026-07-05
- Title:
-
Multimodal graph contrastive learning recommendation based on fusion diffusion models
- Author(s):
-
SHEN Xueli; YU Jiahao
-
School of Software, Liaoning Technical University, Huludao 125105, China
-
- Keywords:
-
recommender system; diffusion model; multimodal; graph contrastive learning; self-supervised learning; multitask learning; data sparsity; user interest
- CLC:
-
TP311
- DOI:
-
10.11992/tis.202507015
- Abstract:
-
To address the inaccurate modeling of user interest preferences caused by noise interference, data sparsity, and cross-modal semantic discrepancies in multimodal data, this paper proposes a diffusion-enhanced exploration framework for multimodal graph contrastive recommendation(DEMGR). First, a diffusion model is employed to generate structurally enhanced contrastive views, effectively suppressing modality-specific noise during self-supervised learning. Second, an isomorphic graph integrating users and items is constructed to strengthen graph connectivity and mitigate the effects of data sparsity on graph representation learning. Finally, a joint optimization framework is designed to combine multitask learning with diffusion enhancement, enabling the collaborative optimization of recommendation objectives and enhancement signals while learning semantically consistent and generalizable multimodal representations. Experiments conducted on three real-world datasets, Baby, Sports, and Clothing, demonstrate that DEMGR consistently outperforms existing mainstream models in terms of Recall and normalized discounted cumulative gain(NDCG), achieving maximum improvements of 5.45% and 5.01%, respectively. These results verify the effectiveness of the proposed method for multimodal recommendation and provide valuable insights for the optimization of multimodal recommender systems.