[1]QIU Jiaqing,DOU Liyun,WANG Jin.Multistage image style transfer based on cross-modal attention[J].CAAI Transactions on Intelligent Systems,2026,21(3):751-762.[doi:10.11992/tis.202508011]
Copy
CAAI Transactions on Intelligent Systems[ISSN 1673-4785/CN 23-1538/TP] Volume:
21
Number of periods:
2026 3
Page number:
751-762
Column:
学术论文—智能系统
Public date:
2026-05-05
- Title:
-
Multistage image style transfer based on cross-modal attention
- Author(s):
-
QIU Jiaqing; DOU Liyun; WANG Jin
-
School of Artificial Intelligence and Computer Science, Nantong University, Nantong 226019, China
-
- Keywords:
-
image style transfer; cross-modal attention; adaptive style modulation; latent diffusion model; image generation; text-guided image generation; multimodal generation; artistic style transfer
- CLC:
-
TP391;TH212
- DOI:
-
10.11992/tis.202508011
- Abstract:
-
Image style transfer (IST) aims to fuse image content with a target artistic style to generate semantically rational and visually expressive images, and it has been widely applied in art creation, personalized image editing, and other fields. Existing methods suffer from low computational efficiency, weak style controllability, loss of details, and disconnection between style and content when handling high-resolution images, complex textures, and cross-modal guided tasks. To address these issues, this paper proposes CAST-Diff, a multistage image style transfer framework built on latent diffusion models. CAST-Diff follows a decoupled collaborative design that combines cross-modal semantic guidance, adaptive regional style regulation, and latent space diffusion refinement so that text and image features can be aligned through cross-modal attention, regional style strength can be adjusted by the adaptive module, and denoising together with detail reconstruction can be completed by the latent diffusion model. Experimental comparisons with mainstream methods such as StyleGAN-Diffusion and ControlNet on the COCO and Flickr30k datasets show that CAST-Diff performs better in style consistency, detail fidelity, and visual naturalness while preserving image structure and fine textures in complex scenes, producing more natural and realistic style transfer results, and improving computational efficiency as well as generalization ability. These advantages make CAST-Diff a practical solution for text-guided high-precision style transfer.