[1]张伟,罗亚远.AI图像生成的精细化可控:结构化与非结构化提示词及其轻量微调的对比研究[J].智能系统学报,2026,21(4):908-918.[doi:10.11992/tis.202509016]
 ZHANG Wei,LUO Yayuan.Fine-tuned control of AI image generation: a comparative study of structured and unstructured prompts and lightweight model fine-tuning[J].CAAI Transactions on Intelligent Systems,2026,21(4):908-918.[doi:10.11992/tis.202509016]
点击复制

AI图像生成的精细化可控:结构化与非结构化提示词及其轻量微调的对比研究

参考文献/References:
[1] XU Junping, ZHANG Xiaolin, LI Hui, et al. Is everyone an artist? a study on user experience of AI-based painting system[J]. Applied sciences, 2023, 13(11): 6496
[2] WU Zhuohao, JI Danwen, YU Kaiwen, et al. AI creativity and the human-AI co-creation model[C]//Human-Computer Interaction. Theory, Methods and Tools. Online: HCII, 2021.
[3] WANG Changsheng. Art innovation or plagiarism? Chinese students’ attitudes toward AI painting technology and influencing factors[J]. IEEE access, 2024, 12: 85795-85805
[4] REDDY A. Artificial everyday creativity: creative leaps with AI through critical making[J]. Digital creativity, 2022, 33(4): 295-313
[5] LIU Jiachang, SHEN Dinghan, ZHANG Yizhe, et al. What makes good in-context examples for GPT-3?[C]//Proceedings of Deep Learning Inside Out: The 3rd Workshop on Knowledge Extraction and Integration for Deep Learning Architectures. Dublin: ACL, 2022.
[6] WITTEVEEN S, ANDREWS M. Investigating prompt engineering in diffusion models[EB/OL]. (2022-11-21)[2025-08-20]. https://arxiv.org/abs/2211.15462.
[7] GOODFELLOW I J, POUGET-ABADIE J, MIRZA M, et al. Generative adversarial nets[J]. Advances in neural information processing systems, 2014, 27: 2672-2680
[8] DENTON E L, CHINTALA S, SZLAM A. Deep generative image models using a laplacian pyramid of adversarial networks[J]. Advances in neural information processing systems, 2015, 28: 1486-1494
[9] 胡铭菲, 左信, 刘建伟. 深度生成模型综述[J]. 自动化学报, 2022, 48(1): 40-74 HU Mingfei, ZUO Xin, LIU Jianwei. Survey on deep generative model[J]. Acta automatica sinica, 2022, 48(1): 40-74
[10] 谢天圻, 吴媛媛, 敬超, 等. GAN模型生成图像检测方法综述[J]. 计算机工程与应用, 2024, 60(22): 74-86 XIE Tianqi, WU Yuanyuan, JING Chao, et al. Survey of image detection methods generated by GAN models[J]. Computer engineering and applications, 2024, 60(22): 74-86
[11] GULRAJANI I, AHMED F, ARJOVSKY M, et al. Improved training of wasserstein gans[J]. Advances in neural information processing systems, 2017, 30: 5767-5777
[12] RADFORD A, METZ L, CHINTALA S. Unsupervised representation learning with deep convolutional generative adversarial networks[EB/OL]. (2015-11-19)[2025-08-20]. https://arxiv.org/abs/1511.06434.
[13] REED S, AKATA Z, YAN X, et al. Generative adversarial text to image synthesis[C]//International Conference on Machine Learning. New York: PMLR, 2016.
[14] ZHANG Han, XU Tao, LI Hongsheng, et al. StackGAN: text to photo-realistic image synthesis with stacked generative adversarial networks[C]//2017 IEEE International Conference on Computer Vision. Venice: IEEE, 2017.
[15] RAMESH A, DHARIWAL P, NICHOL A, et al. Hierarchical text-conditional image generation with clip latents[EB/OL]. (2022-04-13)[2025-08-16]. https://arxiv.org/abs/2204.06125.
[16] CHAN W, DENTON E, FLEET D, et al. Photorealistic text-to-image diffusion models with deep language understanding[C]//Advances in Neural Information Processing Systems 35. New Orleans: NeurIPS, 2022.
[17] RADFORD A, KIM J W, HALLACY C, et al. Learning transferable visual models from natural language supervision[C]//International Conference on Machine Learning. Online: PmLR, 2021.
[18] NICHOL A, DHARIWAL P, RAMESH A, et al. GLIDE: towards photorealistic image generation and editing with text-guided diffusion models[EB/OL]. (2021-12-20)[2025-08-16]. https://arxiv.org/abs/2112.10741.
[19] SHAVLOKHOVA V, VOLLMER A, ZOUBOULIS C C, et al. Finetuning of GLIDE stable diffusion model for AI-based text-conditional image synthesis of dermoscopic images[J]. Frontiers in medicine, 2023, 10: 1231436
[20] AVRAHAMI O, LISCHINSKI D, FRIED O. Blended diffusion for text-driven editing of natural images[C]//2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New Orleans: IEEE, 2022.
[21] KIM G, KWON T, YE J C. DiffusionCLIP: text-guided diffusion models for robust image manipulation[C]//2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New Orleans: IEEE, 2022
[22] AVRAHAMI O, FRIED O, LISCHINSKI D. Blended latent diffusion[J]. ACM transactions on graphics, 2023, 42(4): 1-11
[23] HERTZ A, MOKADY R, TENENBAUM J, et al. Prompt-to-prompt image editing with cross attention control[EB/OL]. (2022-10-02)[2025-08-16]. https://arxiv.org/abs/2208.01626.
[24] OPPENLAENDER J, LINDER R, SILVENNOINEN J. Prompting AI art: an investigation into the creative skill of prompt engineering[J]. International journal of human–computer interaction, 2025, 41(16): 10207-10229
[25] OPPENLAENDER J. The creativity of text-to-image generation[C]//Proceedings of the 25th International Academic Mindtrek Conference. Tampere: ACM, 2022.
[26] BROWNE K. Who (or what) is an AI artist?[J]. Leonardo, 2022, 55(2): 130-134
[27] 王常圣. 面向大模型艺术图像生成的提示词工程研究[J]. 图学学报, 2024, 45(6): 1243-1255 WANG Changsheng. Research on prompt engineering for generating artistic images using large models[J]. Journal of graphic sciences, 2024, 45(6): 1243-1255
[28] DAI Kanyuan, SHAO Ji, GONG Bo, et al. CLIP-FSSC: a transferable visual model for fish and shrimp species classification based on natural language supervision[J]. Aquacultural engineering, 2024, 107: 102460
[29] WANG Jianyi, CHAN K C K, LOY C C. Exploring CLIP for assessing the look and feel of images[J]. Proceedings of the AAAI conference on artificial intelligence, 2023, 37(2): 2555-2563
[30] HENTSCHEL S, KOBS K, HOTHO A. CLIP knows image aesthetics[J]. Frontiers in artificial intelligence, 2022, 5: 976235
[31] 郑明琪, 陈晓慧, 刘冰, 等. 提示学习中思维链生成和增强方法综述[J]. 计算机科学, 2025, 52(1): 56-64 ZHENG Mingqi, CHEN Xiaohui, LIU Bing, et al. A review of methods for generating and enhancing thought chains in prompted learning[J]. Computer science, 2025, 52(1): 56-64
[32] ZHANG Wei, JIN Wei, RHO S, et al. A federated learning framework for brain tumor segmentation without sharing patient data[J]. International journal of imaging systems and technology, 2024, 34(4): e23147

备注/Memo

收稿日期:2025-9-9。
基金项目:Shenzhen Science and Technology Innovation Program(JCYJ20240813143102004).
作者简介:张伟,博士后,主要研究方向为计算机视觉、深度学习。发表学术论文30余篇。E-mail:wei.zhang.2@ki.se。;罗亚远,本科生,主要研究方向为计算机视觉。E-mail:luoyayuan.pl@qq.com。
通讯作者:张伟. E-mail:wei.zhang.2@ki.se

更新日期/Last Update: 1900-01-01
Copyright © 《 智能系统学报》 编辑部
地址:(150001)黑龙江省哈尔滨市南岗区南通大街145-1号楼 电话:0451- 82534001、82518134 邮箱:tis@vip.sina.com