Stable Diffusion 1.5 ControlNet Inpaint加载自定义LoRA无效果求助
问题:LoRA在ControlNet Inpaint流程中未触发的优化方案
背景
我尝试用基于runwayml/stable-diffusion-v1-5微调的LoRA做图像修复,因为基础模型修复效果差,引入了lllyasviel/control_v11p_sd15_inpaint ControlNet模型。另外尝试在runwayml/stable-diffusion-inpainting上训练LoRA时,工具提示维度不匹配报错。
LoRA本身是有效的——用以下代码加载LoRA,触发词sks chair能生成符合预期的目标主体图像:
from diffusers import StableDiffusionPipeline, ControlNetModel, DDIMScheduler from diffusers.utils import load_image import numpy as np import torch from PIL import Image generator = torch.Generator(device="cpu").manual_seed(2) pipe = StableDiffusionPipeline.from_pretrained( "runwayml/stable-diffusion-v1-5", torch_dtype=torch.float16 ) pipe.load_lora_weights('/content/drive/MyDrive/Loras', weight_name="fine_tuned_lora.safetensors") _ = pipe.to("cuda") # generate image image = pipe( "sks chair", guidance_scale=8, num_inference_steps=40, generator=generator, cross_attention_kwargs={"scale": 0.99}, eta=1.0, ).images[0] image
(正确结果为目标椅子图像)
但加入ControlNet Inpaint后,LoRA未被触发,生成结果不符合预期,代码如下:
from diffusers import StableDiffusionPipeline, ControlNetModel, DDIMScheduler, StableDiffusionControlNetInpaintPipeline from diffusers.utils import load_image import numpy as np import torch from PIL import Image init_image = Image.open("empty_room_image.jpg").convert("RGB") mask_image = Image.open("empty_room_image_mask.jpg").convert("RGB") init_image = init_image.resize((512, 512)) mask_image = mask_image.resize((512, 512)) def make_inpaint_condition(image, image_mask): image = np.array(image.convert("RGB")).astype(np.float32) / 255.0 image_mask = np.array(image_mask.convert("L")).astype(np.float32) / 255.0 assert image.shape[0:1] == image_mask.shape[0:1], "image and image_mask must have the same image size" image[image_mask > 0.5] = -1.0 # set as masked pixel image = np.expand_dims(image, 0).transpose(0, 3, 1, 2) image = torch.from_numpy(image) return image control_image = make_inpaint_condition(init_image, mask_image) controlnet = ControlNetModel.from_pretrained( "lllyasviel/control_v11p_sd15_inpaint", torch_dtype=torch.float16 ) pipe = StableDiffusionControlNetInpaintPipeline.from_pretrained( "runwayml/stable-diffusion-v1-5", controlnet=controlnet, torch_dtype=torch.float16 ) pipe.load_lora_weights('/content/drive/MyDrive/Loras', weight_name="fine_tuned_lora.safetensors") _ = pipe.to("cuda") image = pipe( prompt="sks chair", num_inference_steps=20, guidance_scale=8, #around 8 looks good controlnet_conditioning_scale=0.9, control_guidance_end=0.9, cross_attention_kwargs={"scale": .96}, generator=generator, eta=1.0, image=init_image, mask_image=mask_image, control_image=control_image, ).images[0] image
(结果未生成目标椅子)
尝试SDXL标准修复也有类似问题,求优化方案。
优化方案
1. 调整LoRA与ControlNet的权重平衡
- 提高
cross_attention_kwargs={"scale": 1.2}(原0.96),增强LoRA对交叉注意力层的影响;同时降低controlnet_conditioning_scale至0.6-0.7,避免ControlNet的结构约束过强压制LoRA的目标特征。 - 可尝试分阶段权重调整:前10步用较高ControlNet权重保证修复区域的结构合理性,后10步降低ControlNet权重并提高LoRA权重,强化目标主体生成。
2. 修正ControlNet输入格式
- 替换自定义的
make_inpaint_condition函数,使用diffusers内置的prepare_image工具生成符合模型要求的输入,避免格式错误:from diffusers.utils import prepare_image control_image = prepare_image(init_image, mask_image, width=512, height=512) - 移除
control_guidance_end=0.9或设置为0.7,让LoRA在生成后期有更多发挥空间。
3. 更换Inpaint管道提升LoRA兼容性
- 尝试用
StableDiffusionInpaintPipeline结合ControlNet的间接方式,部分场景下这种组合对LoRA的兼容性更好:from diffusers import StableDiffusionInpaintPipeline pipe = StableDiffusionInpaintPipeline.from_pretrained( "runwayml/stable-diffusion-v1-5", torch_dtype=torch.float16 ) pipe.load_lora_weights('/content/drive/MyDrive/Loras', weight_name="fine_tuned_lora.safetensors") # 注入ControlNet pipe.controlnet = controlnet pipe.to("cuda")
4. 解决Inpaint模型训练LoRA的维度问题
runwayml/stable-diffusion-inpainting的UNet结构与SDv1-5略有差异,训练LoRA时需指定匹配的目标层:使用peft或diffusers的LoRA训练工具时,明确指定基础模型为runwayml/stable-diffusion-inpainting,并确保target_modules与该模型的UNet模块名称一致(部分模块后缀含_inpaint)。- 选择支持Inpaint模型的第三方训练模板,避免维度不匹配报错。
5. 强化提示词引导
- 在提示词中提升触发词权重,比如
(sks chair:1.2),同时补充环境描述(如sks chair in empty room, photorealistic),让模型更明确生成目标。
内容的提问来源于stack exchange,提问作者Pieter Bosma
相关产品推荐
相关产品推荐

