[1]张伟,罗亚远.AI图像生成的精细化可控:结构化与非结构化提示词及其轻量微调的对比研究[J].智能系统学报,2026,21(4):908-918.[doi:10.11992/tis.202509016]
ZHANG Wei,LUO Yayuan.Fine-tuned control of AI image generation: a comparative study of structured and unstructured prompts and lightweight model fine-tuning[J].CAAI Transactions on Intelligent Systems,2026,21(4):908-918.[doi:10.11992/tis.202509016]
点击复制
《智能系统学报》[ISSN 1673-4785/CN 23-1538/TP] 卷:
21
期数:
2026年第4期
页码:
908-918
栏目:
学术论文—机器学习
出版日期:
2026-07-05
- Title:
-
Fine-tuned control of AI image generation: a comparative study of structured and unstructured prompts and lightweight model fine-tuning
- 作者:
-
张伟1, 罗亚远2
-
1. 卡罗琳斯卡学院 临床科学系, 斯德哥尔摩 胡丁厄 14152;
2. 辽宁工程技术大学 软件学院, 辽宁 葫芦岛 125000
- Author(s):
-
ZHANG Wei1, LUO Yayuan2
-
1. Department of Clinical Science, Karolinska Institutet, Huddinge 14152, Sweden;
2. School of Software Engineering, Liaoning Technical University, Huludao 125000, China
-
- 关键词:
-
生成式人工智能; 提示词工程; 结构化提示; LoRA模型; CLIPScore; 可控生成; 美学评分; 智能包装设计
- Keywords:
-
artificial intelligence generated content; prompt engineering; structured prompts; LoRA model; CLIPScore; controlled generation; aesthetic scoring; intelligent packaging design
- 分类号:
-
TP391.41; TB482
- DOI:
-
10.11992/tis.202509016
- 摘要:
-
针对文本到图像(T2I)生成中结果不稳定、难以控图的痛点,结合生成式人工智能(artificial intelligence generated content,AIGC)技术中的提示词和微调模型,通过分析不同格式的提示词及其微调模型对控图能力的影响,建立“结构化提示词及LoRA微调”可复现生图模型以解决当前包装设计过程中LLM(large language model)绘图对生成结果的精细化可控等问题。以FLUX-dev为基座,围绕100个主题构建提示词结构化与否、LoRA微调与否的四象限数据集;以CLIPScore与自动化美学评分为基准,配合线性混合效应与配对统计,并辅以同构图的可视化对照与全局—细节双重检验。在此基础上,增加DALL·E3跨模态模型以验证结构化语义一致性的普适性。结果表明:结构化主要稳定全局版式与层级关系,LoRA打磨边缘与材质细节;针对基础文本生图模型,结构化提示词能稳健提升图文符合度;LoRA微调能显著提升美学评分。此外,不同大模型侧重不同,新版本并不必然优于旧版本,匹配符合版本的高质量提示词微调组合尤为关键。本结论可为未来的提示词工程和低成本模型定制提供理论基础和实践参考。
- Abstract:
-
To mitigate instability and poor controllability in text-to-image(T2I) generation, this study integrates prompt engineering and lightweight fine-tuning in artificial-intelligence-generated content(AIGC) systems. We examine how prompt formats and fine-tuned models jointly affect controllability of image generation and propose a reproducible pipeline built on “structured prompts + LoRA fine-tuning.” The method addresses the challenge of achieving precise control in LLM-assisted image generation for packaging design workflows. Building on FLUX-dev, we create a four-quadrant, 100-theme dataset that systematically varies prompt structure and LoRA configurations. The proposed method is evaluated using CLIPScore and automated aesthetic ratings. Statistical significance is examined through linear mixed-effects models and paired tests, while visual quality is assessed through isomorphic renderings and dual global–detail inspections. To assess the generalizability of the structured prompt strategy, we also evaluate the dataset with the cross-modal DALL·E3 model. Results show that structured prompts mainly stabilize overall layout and hierarchy, whereas LoRA chiefly sharpens edges and material details. For base T2I models, structured prompts consistently improve image–text alignment, while LoRA substantially elevates aesthetic scores. Different large models exhibit distinct strengths, and newer versions do not inherently surpass older ones; the key is pairing high-quality prompts with LoRA settings suited to each version. These findings provide both theoretical insights and practical guidance for prompt engineering and cost-effective model adaptation.
更新日期/Last Update:
1900-01-01