The integration of artificial intelligence (AI) and art design has unlocked new potential for personalized, creative, and emotion-driven content generation. However, existing methods still face major challenges in style controllability, emotion consistency, and user satisfaction prediction. This study proposes a multimodal AI art generation strategy based on Stable Diffusion and BERT, using ControlNet, text-to-image (T2I)-Adapter, style–emotion mapping (SEM), sentiment prediction optimization, guided score distillation (GSD), and score-based generative model (SGM) to achieve high-quality, personalized art generation. First, ControlNet and T2I-Adapter are introduced into the Stable Diffusion method to enhance the fine control of text descriptions, visual references, and style labels, and improve the controllability and emotion consistency of generated content. In addition, a SEM model is constructed to establish a deep correspondence between user emotions and visual aesthetics using multimodal feature learning (including color, composition, and style attributes). Afterwards, GSD and SGM are used to optimize the diffusion model to minimize the interference of irrelevant information a
📖 افتح في inklap 🔗 DOI 📮 اطلب بحثاً