训练IP-Adapter plus模型后推理出现RuntimeError问题求助
问题描述
我从腾讯AILab的IP-Adapter仓库下载了代码包,执行以下命令训练文本+图片输入、图片输出的IP-Adapter plus模型:
accelerate launch --num_processes 2 --multi_gpu --mixed_precision "fp16" \ tutorial_train_plus.py \ --pretrained_model_name_or_path="stable-diffusion-v1-5/" \ --image_encoder_path="models/image_encoder/" \ --data_json_file="assets/prompt_image.json" \ --data_root_path="assets/train/" \ --mixed_precision="fp16" \ --resolution=512 \ --train_batch_size=2 \ --dataloader_num_workers=4 \ --learning_rate=1e-04 \ --weight_decay=0.01 \ --output_dir="out_model/" \ --save_steps=3
训练过程中出现以下提示,但训练仍能继续:
保存时移除了共享张量 {'adapter_modules.27.to_k_ip.weight', 'adapter_modules.1.to_v_ip.weight', 'adapter_modules.31.to_k_ip.weight', 'adapter_modules.15.to_k_ip.weight', 'adapter_modules.31.to_v_ip.weight', 'adapter_modules.11.to_k_ip.weight', 'adapter_modules.23.to_k_ip.weight', 'adapter_modules.3.to_k_ip.weight', 'adapter_modules.25.to_v_ip.weight', 'adapter_modules.21.to_k_ip.weight', 'adapter_modules.17.to_v_ip.weight', 'adapter_modules.13.to_k_ip.weight', 'adapter_modules.17.to_k_ip.weight', 'adapter_modules.19.to_v_ip.weight', 'adapter_modules.13.to_v_ip.weight', 'adapter_modules.7.to_v_ip.weight', 'adapter_modules.7.to_k_ip.weight', 'adapter_modules.29.to_k_ip.weight', 'adapter_modules.3.to_v_ip.weight', 'adapter_modules.5.to_v_ip.weight', 'adapter_modules.21.to_v_ip.weight', 'adapter_modules.5.to_k_ip.weight', 'adapter_modules.23.to_v_ip.weight', 'adapter_modules.25.to_k_ip.weight', 'adapter_modules.1.to_k_ip.weight', 'adapter_modules.9.to_v_ip.weight', 'adapter_modules.9.to_k_ip.weight', 'adapter_modules.15.to_v_ip.weight', 'adapter_modules.27.to_v_ip.weight', 'adapter_modules.29.to_v_ip.weight', 'adapter_modules.19.to_k_ip.weight', 'adapter_modules.11.to_v_ip.weight'}。这应该没问题,但可以通过重新加载时是否有警告来验证。
训练完成后将权重转换为ip_adapter.bin,运行推理代码ip_adapter-plus_demo.py,模型路径配置如下:
base_model_path = "SG161222/Realistic_Vision_V4.0_noVAE" vae_model_path = "stabilityai/sd-vae-ft-mse" image_encoder_path = "models/image_encoder" ip_ckpt = "out_model/demo_plus_checkpoint/ip_adapter.bin"
运行时出现如下错误:
RuntimeError: 加载ModuleList的state_dict时出错: state_dict中缺少以下key:"1.to_k_ip.weight", "1.to_v_ip.weight", "3.to_k_ip.weight", "3.to_v_ip.weight", "5.to_k_ip.weight", "5.to_v_ip.weight", "7.to_k_ip.weight", "7.to_v_ip.weight", "9.to_k_ip.weight", "9.to_v_ip.weight", "11.to_k_ip.weight", "11.to_v_ip.weight", "13.to_k_ip.weight", "13.to_v_ip.weight", "15.to_k_ip.weight", "15.to_v_ip.weight", "17.to_k_ip.weight", "17.to_v_ip.weight", "19.to_k_ip.weight", "19.to_v_ip.weight", "21.to_k_ip.weight", "21.to_v_ip.weight", "23.to_k_ip.weight", "23.to_v_ip.weight", "25.to_k_ip.weight", "25.to_v_ip.weight", "27.to_k_ip.weight", "27.to_v_ip.weight", "29.to_k_ip.weight", "29.to_v_ip.weight", "31.to_k_ip.weight", "31.to_v_ip.weight".
请问哪一步操作出错导致了该问题?
问题分析与解决
核心出错点
- 多GPU训练时共享张量未完整保存:训练用2卡多GPU模式,模型的adapter层参数被设为共享张量以节省显存,但保存checkpoint时这些共享张量被自动移除——训练时的提示已经明确指出这一点,只是未被重视,最终导致权重文件缺失关键参数。
- 权重转换环节未补全缺失参数:将训练生成的checkpoint转为
ip_adapter.bin时,未处理被移除的共享张量,转换后的权重文件依然缺少对应参数。 - 训练与推理基础模型不匹配:训练用的是
stable-diffusion-v1-5,但推理用的是Realistic_Vision_V4.0_noVAE,虽同属SDv1架构,但模块命名或细节结构存在差异,进一步加剧了参数加载失败的问题。
修复步骤
- 修改训练保存逻辑:在
tutorial_train_plus.py的模型保存代码段中,添加禁用共享张量自动剔除的配置,比如将保存代码改为:
确保共享张量被完整保存。trainer.save_model(output_dir, safe_serialization=False) - 重新生成权重文件:用修改后的训练脚本重新训练(或直接加载完整的训练checkpoint),再重新转换生成
ip_adapter.bin。 - 统一基础模型:推理时改用训练时的
stable-diffusion-v1-5作为基础模型,避免因模型差异导致的参数映射问题。
内容的提问来源于stack exchange,提问作者weiming
相关产品推荐
相关产品推荐

