[1]ZHANG Wei,LUO Yayuan.Fine-tuned control of AI image generation: a comparative study of structured and unstructured prompts and lightweight model fine-tuning[J].CAAI Transactions on Intelligent Systems,2026,21(4):908-918.[doi:10.11992/tis.202509016]
Copy
CAAI Transactions on Intelligent Systems[ISSN 1673-4785/CN 23-1538/TP] Volume:
21
Number of periods:
2026 4
Page number:
908-918
Column:
学术论文—机器学习
Public date:
2026-07-05
- Title:
-
Fine-tuned control of AI image generation: a comparative study of structured and unstructured prompts and lightweight model fine-tuning
- Author(s):
-
ZHANG Wei1; LUO Yayuan2
-
1. Department of Clinical Science, Karolinska Institutet, Huddinge 14152, Sweden;
2. School of Software Engineering, Liaoning Technical University, Huludao 125000, China
-
- Keywords:
-
artificial intelligence generated content; prompt engineering; structured prompts; LoRA model; CLIPScore; controlled generation; aesthetic scoring; intelligent packaging design
- CLC:
-
TP391.41; TB482
- DOI:
-
10.11992/tis.202509016
- Abstract:
-
To mitigate instability and poor controllability in text-to-image(T2I) generation, this study integrates prompt engineering and lightweight fine-tuning in artificial-intelligence-generated content(AIGC) systems. We examine how prompt formats and fine-tuned models jointly affect controllability of image generation and propose a reproducible pipeline built on “structured prompts + LoRA fine-tuning.” The method addresses the challenge of achieving precise control in LLM-assisted image generation for packaging design workflows. Building on FLUX-dev, we create a four-quadrant, 100-theme dataset that systematically varies prompt structure and LoRA configurations. The proposed method is evaluated using CLIPScore and automated aesthetic ratings. Statistical significance is examined through linear mixed-effects models and paired tests, while visual quality is assessed through isomorphic renderings and dual global–detail inspections. To assess the generalizability of the structured prompt strategy, we also evaluate the dataset with the cross-modal DALL·E3 model. Results show that structured prompts mainly stabilize overall layout and hierarchy, whereas LoRA chiefly sharpens edges and material details. For base T2I models, structured prompts consistently improve image–text alignment, while LoRA substantially elevates aesthetic scores. Different large models exhibit distinct strengths, and newer versions do not inherently surpass older ones; the key is pairing high-quality prompts with LoRA settings suited to each version. These findings provide both theoretical insights and practical guidance for prompt engineering and cost-effective model adaptation.