You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

训练IP-Adapter plus模型后推理出现RuntimeError问题求助

问题描述

我从腾讯AILab的IP-Adapter仓库下载了代码包,执行以下命令训练文本+图片输入、图片输出的IP-Adapter plus模型:

accelerate launch --num_processes 2 --multi_gpu --mixed_precision "fp16" \
  tutorial_train_plus.py \
  --pretrained_model_name_or_path="stable-diffusion-v1-5/" \
  --image_encoder_path="models/image_encoder/" \
  --data_json_file="assets/prompt_image.json" \
  --data_root_path="assets/train/" \
  --mixed_precision="fp16" \
  --resolution=512 \
  --train_batch_size=2 \
  --dataloader_num_workers=4 \
  --learning_rate=1e-04 \
  --weight_decay=0.01 \
  --output_dir="out_model/" \
  --save_steps=3

训练过程中出现以下提示,但训练仍能继续:

保存时移除了共享张量 {'adapter_modules.27.to_k_ip.weight', 'adapter_modules.1.to_v_ip.weight', 'adapter_modules.31.to_k_ip.weight', 'adapter_modules.15.to_k_ip.weight', 'adapter_modules.31.to_v_ip.weight', 'adapter_modules.11.to_k_ip.weight', 'adapter_modules.23.to_k_ip.weight', 'adapter_modules.3.to_k_ip.weight', 'adapter_modules.25.to_v_ip.weight', 'adapter_modules.21.to_k_ip.weight', 'adapter_modules.17.to_v_ip.weight', 'adapter_modules.13.to_k_ip.weight', 'adapter_modules.17.to_k_ip.weight', 'adapter_modules.19.to_v_ip.weight', 'adapter_modules.13.to_v_ip.weight', 'adapter_modules.7.to_v_ip.weight', 'adapter_modules.7.to_k_ip.weight', 'adapter_modules.29.to_k_ip.weight', 'adapter_modules.3.to_v_ip.weight', 'adapter_modules.5.to_v_ip.weight', 'adapter_modules.21.to_v_ip.weight', 'adapter_modules.5.to_k_ip.weight', 'adapter_modules.23.to_v_ip.weight', 'adapter_modules.25.to_k_ip.weight', 'adapter_modules.1.to_k_ip.weight', 'adapter_modules.9.to_v_ip.weight', 'adapter_modules.9.to_k_ip.weight', 'adapter_modules.15.to_v_ip.weight', 'adapter_modules.27.to_v_ip.weight', 'adapter_modules.29.to_v_ip.weight', 'adapter_modules.19.to_k_ip.weight', 'adapter_modules.11.to_v_ip.weight'}。这应该没问题,但可以通过重新加载时是否有警告来验证。

训练完成后将权重转换为ip_adapter.bin,运行推理代码ip_adapter-plus_demo.py,模型路径配置如下:

base_model_path = "SG161222/Realistic_Vision_V4.0_noVAE"
vae_model_path = "stabilityai/sd-vae-ft-mse"
image_encoder_path = "models/image_encoder"
ip_ckpt = "out_model/demo_plus_checkpoint/ip_adapter.bin"

运行时出现如下错误:

RuntimeError: 加载ModuleList的state_dict时出错:
        state_dict中缺少以下key:"1.to_k_ip.weight", "1.to_v_ip.weight", "3.to_k_ip.weight", "3.to_v_ip.weight", "5.to_k_ip.weight", "5.to_v_ip.weight", "7.to_k_ip.weight", "7.to_v_ip.weight", "9.to_k_ip.weight", "9.to_v_ip.weight", "11.to_k_ip.weight", "11.to_v_ip.weight", "13.to_k_ip.weight", "13.to_v_ip.weight", "15.to_k_ip.weight", "15.to_v_ip.weight", "17.to_k_ip.weight", "17.to_v_ip.weight", "19.to_k_ip.weight", "19.to_v_ip.weight", "21.to_k_ip.weight", "21.to_v_ip.weight", "23.to_k_ip.weight", "23.to_v_ip.weight", "25.to_k_ip.weight", "25.to_v_ip.weight", "27.to_k_ip.weight", "27.to_v_ip.weight", "29.to_k_ip.weight", "29.to_v_ip.weight", "31.to_k_ip.weight", "31.to_v_ip.weight".

请问哪一步操作出错导致了该问题?


问题分析与解决

核心出错点

  1. 多GPU训练时共享张量未完整保存:训练用2卡多GPU模式,模型的adapter层参数被设为共享张量以节省显存,但保存checkpoint时这些共享张量被自动移除——训练时的提示已经明确指出这一点,只是未被重视,最终导致权重文件缺失关键参数。
  2. 权重转换环节未补全缺失参数:将训练生成的checkpoint转为ip_adapter.bin时,未处理被移除的共享张量,转换后的权重文件依然缺少对应参数。
  3. 训练与推理基础模型不匹配:训练用的是stable-diffusion-v1-5,但推理用的是Realistic_Vision_V4.0_noVAE,虽同属SDv1架构,但模块命名或细节结构存在差异,进一步加剧了参数加载失败的问题。

修复步骤

  • 修改训练保存逻辑:在tutorial_train_plus.py的模型保存代码段中,添加禁用共享张量自动剔除的配置,比如将保存代码改为:
    trainer.save_model(output_dir, safe_serialization=False)
    
    确保共享张量被完整保存。
  • 重新生成权重文件:用修改后的训练脚本重新训练(或直接加载完整的训练checkpoint),再重新转换生成ip_adapter.bin。
  • 统一基础模型:推理时改用训练时的stable-diffusion-v1-5作为基础模型,避免因模型差异导致的参数映射问题。

内容的提问来源于stack exchange,提问作者weiming

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 18:07:02