stabilityai/stable-diffusion-xl-refiner-1.0模型未正确响应输入图像问题咨询
解决Stable Diffusion XL Refiner不参考输入图像的问题
明确Refiner的使用逻辑:SDXL Refiner是后期精修模型,不能直接用外部图像+提示词生成内容。它的正确流程是先通过SDXL Base模型生成初始图像,再将Base的输出作为Refiner的输入,搭配提示词做细节或风格调整。直接喂外部图像的话,模型会默认忽略图像,仅依据提示词生成。
检查参数设置是否到位:用Diffusers库调用时,必须将输入图像传入
image参数,同时设置合理的strength值(控制修改幅度,一般取0.3-0.7,值过高会大幅偏离原图)。参考代码如下:from diffusers import StableDiffusionXLImg2ImgPipeline import torch from PIL import Image # 加载Refiner模型 refiner = StableDiffusionXLImg2ImgPipeline.from_pretrained( "stabilityai/stable-diffusion-xl-refiner-1.0", torch_dtype=torch.float16 ).to("cuda") # 加载并预处理输入图像(需为RGB格式的PIL图像,尺寸匹配1024x1024) init_image = Image.open("your_input_image.png").convert("RGB").resize((1024, 1024)) prompt = "make it anthropomorphic dog as a doctor" # 生成精修后的图像 refined_image = refiner( prompt=prompt, image=init_image, strength=0.5, guidance_scale=7.5 ).images[0] refined_image.save("final_image.png")验证输入图像规格:Refiner默认适配1024x1024的RGB图像,尺寸不符或格式错误时,模型可能无法识别输入图像,转而仅依赖提示词生成。先将图像调整到对应尺寸、转成RGB格式后再尝试。
优化提示词表述:原提示词中的"it"指代模糊,若输入图像主体不是狗,模型无法建立关联。可以改成更明确的表述,比如输入是普通狗时,换成"turn this dog into an anthropomorphic doctor",让模型清楚要基于输入图像做修改。
内容的提问来源于stack exchange,提问作者Aditya Jhaveri
相关产品推荐
相关产品推荐

