You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用StableDiffusionPipeline生成图像后如何清理GPU内存?

解决StableDiffusionPipeline生成图像后GPU内存未释放的问题

你遇到的GPU内存未释放问题,核心原因有两个:一是每次调用create函数都会重新加载整个模型到GPU,旧模型实例未被及时回收;二是推理后没有显式清理GPU缓存和无用对象。下面是具体的解决办法:

1. 避免重复加载模型(最有效)

当前代码每次调用create都会重新加载Stable Diffusion模型,这不仅耗时,还会导致多个模型实例占用GPU内存。把模型初始化逻辑移到函数外面,只执行一次:

import torch
from pathlib import Path
import time
import calendar
from diffusers import StableDiffusionPipeline, EulerDiscreteScheduler

# 全局初始化模型,仅执行一次
model_id = "SG161222/Realistic_Vision_V2.0"
device = "cuda" if torch.cuda.is_available() else "cpu"
scheduler = EulerDiscreteScheduler.from_pretrained(model_id, subfolder="scheduler")
pipe = StableDiffusionPipeline.from_pretrained(model_id, torch_dtype=torch.float16)
pipe = pipe.to(device)

# 可选:启用内存优化,进一步降低显存占用
pipe.enable_model_cpu_offload()  # 按需将模型层加载到GPU,用完放回CPU

def create(prompt, negative_prompt, steps, scale, num_images_per_prompt, seed, width, height):
    DIR_NAME="../output/"
    dirpath = Path(DIR_NAME)
    dirpath.mkdir(parents=True, exist_ok=True)

    prompt_ = prompt
    negative_prompt_ = negative_prompt
    steps_ = int(steps)
    scale_ = float(scale)
    width_ = int(width)
    height_ = int(height)
    num_images_per_prompt_ = int(num_images_per_prompt)
    seed_ = int(seed, 16)
    if seed_ == -1:
        seed_ = torch.randint(0, 1000000, (1,)).item()

    generator = torch.Generator(device=device).manual_seed(seed_)

    output = pipe(prompt_, negative_prompt=negative_prompt_, width=width_, height=height_, num_inference_steps=steps_,
                guidance_scale=scale_, num_images_per_prompt=num_images_per_prompt_, generator=generator)

    for idx, image in enumerate(output.images):
        current_GMT = time.gmtime()
        time_stamp = calendar.timegm(current_GMT)
        image_name = f"{time_stamp} - {idx}.png"
        image_path = dirpath / image_name
        image.save(image_path)
        print(idx)

2. 显式清理GPU内存(适用于必须动态加载模型的场景)

如果你的业务需要每次调用函数都切换不同模型,那在生成完成后要手动清理无用对象和GPU缓存:

在create函数末尾添加以下代码:

# 删除不再使用的对象
del pipe, scheduler, output
# 强制触发Python垃圾回收
import gc
gc.collect()
# 清空GPU缓存
if torch.cuda.is_available():
    torch.cuda.empty_cache()
    torch.cuda.ipc_collect()

3. 启用diffusers内置内存优化

diffusers提供了多种内存优化工具,能自动管理模型在GPU/CPU之间的调度,减少显存占用的同时,也能让内存更容易被释放:

  • pipe.enable_model_cpu_offload():适合显存较小的设备,模型层按需加载到GPU
  • pipe.enable_sequential_cpu_offload():轻量化的逐块加载模式
  • pipe.enable_attention_slicing():对注意力层进行切片,降低峰值内存占用

这些方法只需在pipe.to(device)之后调用即可,无需手动管理内存。

内容的提问来源于stack exchange,提问作者M. Özdemir

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 00:14:56