You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

StableCascadeCombinedPipeline传入图片时出现张量设备不匹配运行时错误

解决StableCascadeCombinedPipeline图片输入的设备不匹配错误

问题原因

你遇到的Runtime Error: StableCascadeCombinedPipeline: Expected all tensors to be on the same device错误,本质是模型已部署在CUDA设备,但传入的PIL图片是CPU端对象,未被转换成模型所需的CUDA张量,导致设备不匹配。仅用文本生成时无此问题,是因为文本输入的预处理会自动适配模型设备,而图片输入需要手动处理。

解决方案

需要对传入的PIL图片做两步核心处理:

  • 将PIL图片转换成符合模型要求的张量格式
  • 将张量移至模型所在的CUDA设备

具体修改如下:

1. 导入图片预处理工具

添加AutoImageProcessor导入,用于处理图片输入的标准化转换:

from transformers import StableCascadeCombinedPipeline, AutoImageProcessor

2. 初始化图片处理器

和模型同步初始化,确保预处理参数匹配模型要求:

# Constants
repo = "stabilityai/stable-cascade"

# Ensure model and scheduler are initialized in GPU-enabled function
if torch.cuda.is_available():
    pipe = StableCascadeCombinedPipeline.from_pretrained(repo, variant="bf16", torch_dtype=torch.bfloat16)
    pipe.to("cuda")
    # 初始化适配模型的图片处理器
    image_processor = AutoImageProcessor.from_pretrained(repo)

3. 修改生成函数,处理图片输入

恢复图片参数,添加预处理逻辑,将图片转换为CUDA张量:

# The generate function
@spaces.GPU(enable_queue=True)
def generate_image(prompt, images):  
    seed = random.randint(-100000, 100000)
    
    # 预处理图片:自动完成尺寸调整、归一化,转换为CUDA张量并匹配模型 dtype
    processed_images = image_processor(images, return_tensors="pt").to("cuda", torch.bfloat16)
    
    results = pipe(
        prompt=prompt,
        images=processed_images["pixel_values"],  # 传入预处理后的标准化张量
        height=1024,
        width=1024,
        num_inference_steps=20, 
        generator=torch.Generator(device="cuda").manual_seed(seed)
    )
    return results.images[0]

关键说明

  • AutoImageProcessor会自动完成模型要求的图片尺寸调整、像素值归一化等步骤,无需手动处理
  • 预处理后的张量必须和模型使用相同的torch_dtype(此处为bfloat16),并明确移至CUDA设备
  • 若传入多张图片,image_processor会自动批量处理,直接传入即可

内容的提问来源于stack exchange,提问作者Mike Ellis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 01:05:11