You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Stable Diffusion 1.5 ControlNet Inpaint加载自定义LoRA无效果求助

问题:LoRA在ControlNet Inpaint流程中未触发的优化方案

背景

我尝试用基于runwayml/stable-diffusion-v1-5微调的LoRA做图像修复,因为基础模型修复效果差,引入了lllyasviel/control_v11p_sd15_inpaint ControlNet模型。另外尝试在runwayml/stable-diffusion-inpainting上训练LoRA时,工具提示维度不匹配报错。

LoRA本身是有效的——用以下代码加载LoRA,触发词sks chair能生成符合预期的目标主体图像:

from diffusers import StableDiffusionPipeline, ControlNetModel, DDIMScheduler
from diffusers.utils import load_image
import numpy as np
import torch
from PIL import Image

generator = torch.Generator(device="cpu").manual_seed(2)

pipe = StableDiffusionPipeline.from_pretrained(
    "runwayml/stable-diffusion-v1-5", torch_dtype=torch.float16
)
pipe.load_lora_weights('/content/drive/MyDrive/Loras', weight_name="fine_tuned_lora.safetensors")
_ = pipe.to("cuda")

# generate image
image = pipe(
    "sks chair",
    guidance_scale=8,
    num_inference_steps=40,
    generator=generator,
    cross_attention_kwargs={"scale": 0.99},
    eta=1.0,
).images[0]

image

(正确结果为目标椅子图像)

但加入ControlNet Inpaint后,LoRA未被触发,生成结果不符合预期,代码如下:

from diffusers import StableDiffusionPipeline, ControlNetModel, DDIMScheduler, StableDiffusionControlNetInpaintPipeline
from diffusers.utils import load_image
import numpy as np
import torch
from PIL import Image

init_image = Image.open("empty_room_image.jpg").convert("RGB")
mask_image = Image.open("empty_room_image_mask.jpg").convert("RGB")

init_image = init_image.resize((512, 512))
mask_image = mask_image.resize((512, 512))

def make_inpaint_condition(image, image_mask):
    image = np.array(image.convert("RGB")).astype(np.float32) / 255.0
    image_mask = np.array(image_mask.convert("L")).astype(np.float32) / 255.0

    assert image.shape[0:1] == image_mask.shape[0:1], "image and image_mask must have the same image size"
    image[image_mask > 0.5] = -1.0  # set as masked pixel
    image = np.expand_dims(image, 0).transpose(0, 3, 1, 2)
    image = torch.from_numpy(image)
    return image

control_image = make_inpaint_condition(init_image, mask_image)

controlnet = ControlNetModel.from_pretrained(
    "lllyasviel/control_v11p_sd15_inpaint", torch_dtype=torch.float16
)

pipe = StableDiffusionControlNetInpaintPipeline.from_pretrained(
    "runwayml/stable-diffusion-v1-5", controlnet=controlnet, torch_dtype=torch.float16
)
pipe.load_lora_weights('/content/drive/MyDrive/Loras', weight_name="fine_tuned_lora.safetensors")
_ = pipe.to("cuda")

image = pipe(
      prompt="sks chair",
      num_inference_steps=20,
      guidance_scale=8, #around 8 looks good
      controlnet_conditioning_scale=0.9,
      control_guidance_end=0.9,
      cross_attention_kwargs={"scale": .96},
      generator=generator,
      eta=1.0,
      image=init_image,
      mask_image=mask_image,
      control_image=control_image,
  ).images[0]
image

(结果未生成目标椅子)

尝试SDXL标准修复也有类似问题,求优化方案。

优化方案

1. 调整LoRA与ControlNet的权重平衡

  • 提高cross_attention_kwargs={"scale": 1.2}(原0.96),增强LoRA对交叉注意力层的影响;同时降低controlnet_conditioning_scale至0.6-0.7,避免ControlNet的结构约束过强压制LoRA的目标特征。
  • 可尝试分阶段权重调整:前10步用较高ControlNet权重保证修复区域的结构合理性,后10步降低ControlNet权重并提高LoRA权重,强化目标主体生成。

2. 修正ControlNet输入格式

  • 替换自定义的make_inpaint_condition函数,使用diffusers内置的prepare_image工具生成符合模型要求的输入,避免格式错误:
    from diffusers.utils import prepare_image
    control_image = prepare_image(init_image, mask_image, width=512, height=512)
    
  • 移除control_guidance_end=0.9或设置为0.7,让LoRA在生成后期有更多发挥空间。

3. 更换Inpaint管道提升LoRA兼容性

  • 尝试用StableDiffusionInpaintPipeline结合ControlNet的间接方式,部分场景下这种组合对LoRA的兼容性更好:
    from diffusers import StableDiffusionInpaintPipeline
    pipe = StableDiffusionInpaintPipeline.from_pretrained(
        "runwayml/stable-diffusion-v1-5", torch_dtype=torch.float16
    )
    pipe.load_lora_weights('/content/drive/MyDrive/Loras', weight_name="fine_tuned_lora.safetensors")
    # 注入ControlNet
    pipe.controlnet = controlnet
    pipe.to("cuda")
    

4. 解决Inpaint模型训练LoRA的维度问题

  • runwayml/stable-diffusion-inpainting的UNet结构与SDv1-5略有差异,训练LoRA时需指定匹配的目标层:使用peft或diffusers的LoRA训练工具时,明确指定基础模型为runwayml/stable-diffusion-inpainting,并确保target_modules与该模型的UNet模块名称一致(部分模块后缀含_inpaint)。
  • 选择支持Inpaint模型的第三方训练模板,避免维度不匹配报错。

5. 强化提示词引导

  • 在提示词中提升触发词权重,比如(sks chair:1.2),同时补充环境描述(如sks chair in empty room, photorealistic),让模型更明确生成目标。

内容的提问来源于stack exchange,提问作者Pieter Bosma

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 12:32:41