Colab中Blip2图像字幕模型出现RuntimeError形状不匹配问题求助
问题解决:Blip2图像字幕生成RuntimeError(形状不匹配)
错误原因分析
你遇到的RuntimeError: shape mismatch: value tensor of shape [81920] cannot be broadcast to indexing result of shape [0],核心问题是调用model.generate时仅传入图像特征,未提供包含图像占位符token的文本输入:
- Blip2的生成逻辑需要将图像编码后的嵌入,替换到语言模型输入中的特定图像占位符位置
- 当仅传入图像时,
input_ids不存在,导致代码中(input_ids == self.config.image_token_index)的匹配结果为空(形状为[0]),但图像编码后的向量长度为81920,无法完成赋值操作
81920的计算逻辑
这个数值由Blip2-opt-2.7b的参数决定:
- 配套的OPT-2.7b语言模型隐藏层维度
hidden_size=2560 - 图像编码器输出的序列长度为32
- 两者相乘:
32 * 2560 = 81920,即图像编码向量展平后的总元素数
解决方案
方案1:修正输入构造,添加空文本prompt
修改生成函数的输入预处理步骤,让处理器同时处理图像和空文本,自动生成包含图像占位符的input_ids:
def generate_caption(processor, model, image_path): image = PILImage.open(image_path).convert("RGB") print("image shape:", image.size) # 修正原代码语法错误:字符串无法直接拼接元组 device = "cuda" if torch.cuda.is_available() else "cpu" # 关键修改:添加text="",触发处理器生成图像占位符token inputs = processor(images=image, text="", return_tensors="pt").to(device) print("Input shape:", inputs['pixel_values'].shape) print("Device:", device) for key, value in inputs.items(): print(f"Key: {key}, Shape: {value.shape}") with torch.no_grad(): generated_ids = model.generate(**inputs) caption = processor.decode(generated_ids[0], skip_special_tokens=True) return caption
方案2:锁定transformers稳定版本
若怀疑是包更新引发的兼容性问题,可指定安装11月8日前的稳定版本(如4.35.2),修改安装命令:
!pip install -q git+https://github.com/huggingface/peft.git transformers==4.35.2 bitsandbytes datasets
方案3:手动构造含图像token的输入
如果需要更精细的控制,可手动生成包含图像占位符的input_ids:
# 在生成函数中添加以下代码,替换原inputs构造逻辑 image_token = processor.tokenizer.convert_tokens_to_ids(processor.tokenizer.image_token) input_ids = torch.tensor([[image_token]]).to(device) # 构造图像输入 image_inputs = processor(images=image, return_tensors="pt").to(device) # 合并输入 inputs = {**image_inputs, "input_ids": input_ids}
修正调用代码语法错误
原调用代码最后一行缺少右括号,修正后:
image_path = "my_image_path.jpg" caption = generate_caption(processor, model, image_path) print(f"{image_path}: {caption}")
内容的提问来源于stack exchange,提问作者Soroush Hosseinpour
相关产品推荐
相关产品推荐

