SD 1.4 VAE微调报错:权重与输入通道维度不匹配
问题分析与解决方案
一、报错原因排查
错误信息:RuntimeError: Expected weight to be a vector of size equal to the number of channels in input, but got weight of shape [128] and input of shape [128, 1024, 1024]
核心问题是张量维度不匹配,具体触发点在训练循环的输入处理逻辑:
你的DataLoader返回的batch是形状为[batch_size, 3, 1024, 1024]的张量(批次图片),但代码中错误地使用batch[0]提取了单张图片,导致输入丢失了batch维度。在多GPU分布式环境下,这种无batch维度的输入会被模型误解为通道维度异常的张量,最终引发维度不匹配错误。
二、代码修复方案
1. 修正输入处理逻辑
将训练循环中的这行代码:
target = batch[0].to(next(vae.parameters()).dtype)
修改为:
target = batch.to(next(vae.parameters()).dtype)
2. 修复重复梯度清零问题
代码中在optimizer.step()前后重复调用了optimizer.zero_grad(),会导致梯度被提前清零,影响反向传播效果。保留一次即可:
optimizer.zero_grad() accelerator.backward(loss) optimizer.step()
3. 优化检查点保存逻辑
使用accelerator.save时无需手动提取model.module,accelerator会自动处理分布式/单GPU环境下的模型状态字典:
accelerator.save({ "epoch": epoch, "model_state_dict": vae.state_dict(), "optimizer_state_dict": optimizer.state_dict(), "loss": loss, }, checkpoint_path)
三、更简便的VAE微调方法
无需手动编写完整训练循环,可利用Hugging Face生态的工具简化流程:
方法1:使用diffusers.Trainer(推荐)
借助预定义的训练器自动处理分布式、日志、 checkpoint 等逻辑:
import yaml from datasets import load_dataset from diffusers import AutoencoderKL, TrainingArguments, Trainer import torch.nn.functional as F # 加载配置 with open('config.yaml', 'r') as file: config = yaml.safe_load(file) # 加载本地图片数据集 dataset = load_dataset("imagefolder", data_dir=config['dataset']['root_dir']) # 数据预处理 def preprocess(examples): from torchvision.transforms import Compose, Resize, ToTensor, Normalize transform = Compose([ Resize((config['dataset']['image_size'], config['dataset']['image_size'])), ToTensor(), Normalize(mean=[0.5, 0.5, 0.5], std=[0.5, 0.5, 0.5]) ]) examples["pixel_values"] = [transform(img.convert("RGB")) for img in examples["image"]] return examples dataset = dataset.with_transform(preprocess) # 初始化VAE vae = AutoencoderKL.from_pretrained(config['model']['path']) # 定义训练参数 training_args = TrainingArguments( output_dir="./vae-finetuned", per_device_train_batch_size=config['training']['batch_size'], num_train_epochs=config['training']['num_epochs'], learning_rate=config['training']['learning_rate'], gradient_accumulation_steps=config['training']['gradient_accumulation_steps'], logging_dir=config['logging']['tensorboard_dir'], save_steps=10, fp16=True, remove_unused_columns=False, ) # 自定义损失计算 def compute_loss(model, inputs): pixel_values = inputs["pixel_values"].to(model.dtype) posterior = model.encode(pixel_values).latent_dist z = posterior.mode() pred = model.decode(z).sample kl_loss = posterior.kl().mean() mse_loss = F.mse_loss(pred, pixel_values, reduction="mean") return mse_loss + config['training']["kl_scale"] * kl_loss # 启动训练 trainer = Trainer( model=vae, args=training_args, train_dataset=dataset["train"], compute_loss=compute_loss, ) trainer.train() # 保存最终模型 trainer.save_model("./vae-finetuned-final")
方法2:复用官方训练脚本
直接使用diffusers官方提供的VAE微调脚本,通过命令行参数配置训练细节,完全避免自定义循环的麻烦:
accelerate launch --mixed_precision=fp16 train_vae.py \ --dataset_name="imagefolder" \ --data_dir="segmented" \ --output_dir="./vae-finetuned" \ --model_name_or_path="vae1dot4" \ --resolution=1024 \ --batch_size=8 \ --learning_rate=5e-4 \ --num_train_epochs=10 \ --save_steps=10
内容的提问来源于stack exchange,提问作者Norhther
相关产品推荐
相关产品推荐

