You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Stable Diffusion生成的图像传入图像字幕Pipeline?

问题解决:将SDXL生成的图像传入图像字幕Pipeline

问题根源

你的代码中,refiner生成图像后调用了.show()方法,而这个方法的返回值是None,导致image变量被覆盖为None,后续传入图像字幕Pipeline时自然会报错。

修正后的完整代码

## Loading Diffusion Pipeline
from diffusers import DiffusionPipeline
import torch
from transformers import pipeline

## Image Generating
# load both base & refiner
base = DiffusionPipeline.from_pretrained(
    "stabilityai/stable-diffusion-xl-base-1.0", torch_dtype=torch.float32, variant="fp16", use_safetensors=True
)
# 启用CPU offload释放内存
base.enable_model_cpu_offload()

refiner = DiffusionPipeline.from_pretrained(
    "stabilityai/stable-diffusion-xl-refiner-1.0",
    text_encoder_2=base.text_encoder_2,
    vae=base.vae,
    torch_dtype=torch.float32,
    use_safetensors=True,
    variant="fp16",
)
refiner.enable_model_cpu_offload()

# 定义推理步数与分工比例
n_steps = 40
high_noise_frac = 0.8

prompt = "A man playing soccer in a stadium"

# 运行base模型生成 latent 图像
latent_image = base(
    prompt=prompt,
    num_inference_steps=n_steps,
    denoising_end=high_noise_frac,
    output_type="latent",
).images
# 先获取最终生成的PIL图像对象,再单独调用show()显示
image = refiner(
    prompt=prompt,
    num_inference_steps=n_steps,
    denoising_start=high_noise_frac,
    image=latent_image,
).images[0]
image.show()

## Image Captioning
image_to_text = pipeline("image-to-text", model="nlpconnect/vit-gpt2-image-captioning")
# 传入PIL图像对象生成字幕
caption_result = image_to_text(image)
print(caption_result)

关键修改点

  • 分离图像获取与显示:将image = refiner(...).images[0].show()拆分为两步,确保image始终保留有效的PIL图像对象,而非.show()返回的None。
  • 修复内存优化调用:原代码中enable_model_cpu_offload未加括号,属于未执行方法,修正后可有效降低显存占用。
  • 变量名优化:将base模型输出命名为latent_image,和最终生成的图像区分开,代码逻辑更清晰。

内容的提问来源于stack exchange,提问作者fabone

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 09:53:21