You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Colab中Blip2图像字幕模型出现RuntimeError形状不匹配问题求助

问题解决:Blip2图像字幕生成RuntimeError(形状不匹配)

错误原因分析

你遇到的RuntimeError: shape mismatch: value tensor of shape [81920] cannot be broadcast to indexing result of shape [0],核心问题是调用model.generate时仅传入图像特征,未提供包含图像占位符token的文本输入:

  • Blip2的生成逻辑需要将图像编码后的嵌入,替换到语言模型输入中的特定图像占位符位置
  • 当仅传入图像时,input_ids不存在,导致代码中(input_ids == self.config.image_token_index)的匹配结果为空(形状为[0]),但图像编码后的向量长度为81920,无法完成赋值操作

81920的计算逻辑

这个数值由Blip2-opt-2.7b的参数决定:

  • 配套的OPT-2.7b语言模型隐藏层维度hidden_size=2560
  • 图像编码器输出的序列长度为32
  • 两者相乘:32 * 2560 = 81920,即图像编码向量展平后的总元素数

解决方案

方案1:修正输入构造,添加空文本prompt

修改生成函数的输入预处理步骤,让处理器同时处理图像和空文本,自动生成包含图像占位符的input_ids:

def generate_caption(processor, model, image_path):
  image = PILImage.open(image_path).convert("RGB")
  print("image shape:", image.size)  # 修正原代码语法错误:字符串无法直接拼接元组

  device = "cuda" if torch.cuda.is_available() else "cpu"

  # 关键修改:添加text="",触发处理器生成图像占位符token
  inputs = processor(images=image, text="", return_tensors="pt").to(device)

  print("Input shape:", inputs['pixel_values'].shape)
  print("Device:", device)
  for key, value in inputs.items():
    print(f"Key: {key}, Shape: {value.shape}")

  with torch.no_grad():
      generated_ids = model.generate(**inputs)
      caption = processor.decode(generated_ids[0], skip_special_tokens=True)

  return caption

方案2:锁定transformers稳定版本

若怀疑是包更新引发的兼容性问题,可指定安装11月8日前的稳定版本(如4.35.2),修改安装命令:

!pip install -q git+https://github.com/huggingface/peft.git transformers==4.35.2 bitsandbytes datasets

方案3:手动构造含图像token的输入

如果需要更精细的控制,可手动生成包含图像占位符的input_ids:

# 在生成函数中添加以下代码,替换原inputs构造逻辑
image_token = processor.tokenizer.convert_tokens_to_ids(processor.tokenizer.image_token)
input_ids = torch.tensor([[image_token]]).to(device)
# 构造图像输入
image_inputs = processor(images=image, return_tensors="pt").to(device)
# 合并输入
inputs = {**image_inputs, "input_ids": input_ids}

修正调用代码语法错误

原调用代码最后一行缺少右括号,修正后:

image_path = "my_image_path.jpg"
caption = generate_caption(processor, model, image_path)
print(f"{image_path}: {caption}")

内容的提问来源于stack exchange,提问作者Soroush Hosseinpour

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 02:35:13