使用StableDiffusionPipeline生成图像后如何清理GPU内存?
解决StableDiffusionPipeline生成图像后GPU内存未释放的问题
你遇到的GPU内存未释放问题,核心原因有两个:一是每次调用create函数都会重新加载整个模型到GPU,旧模型实例未被及时回收;二是推理后没有显式清理GPU缓存和无用对象。下面是具体的解决办法:
1. 避免重复加载模型(最有效)
当前代码每次调用create都会重新加载Stable Diffusion模型,这不仅耗时,还会导致多个模型实例占用GPU内存。把模型初始化逻辑移到函数外面,只执行一次:
import torch from pathlib import Path import time import calendar from diffusers import StableDiffusionPipeline, EulerDiscreteScheduler # 全局初始化模型,仅执行一次 model_id = "SG161222/Realistic_Vision_V2.0" device = "cuda" if torch.cuda.is_available() else "cpu" scheduler = EulerDiscreteScheduler.from_pretrained(model_id, subfolder="scheduler") pipe = StableDiffusionPipeline.from_pretrained(model_id, torch_dtype=torch.float16) pipe = pipe.to(device) # 可选:启用内存优化,进一步降低显存占用 pipe.enable_model_cpu_offload() # 按需将模型层加载到GPU,用完放回CPU def create(prompt, negative_prompt, steps, scale, num_images_per_prompt, seed, width, height): DIR_NAME="../output/" dirpath = Path(DIR_NAME) dirpath.mkdir(parents=True, exist_ok=True) prompt_ = prompt negative_prompt_ = negative_prompt steps_ = int(steps) scale_ = float(scale) width_ = int(width) height_ = int(height) num_images_per_prompt_ = int(num_images_per_prompt) seed_ = int(seed, 16) if seed_ == -1: seed_ = torch.randint(0, 1000000, (1,)).item() generator = torch.Generator(device=device).manual_seed(seed_) output = pipe(prompt_, negative_prompt=negative_prompt_, width=width_, height=height_, num_inference_steps=steps_, guidance_scale=scale_, num_images_per_prompt=num_images_per_prompt_, generator=generator) for idx, image in enumerate(output.images): current_GMT = time.gmtime() time_stamp = calendar.timegm(current_GMT) image_name = f"{time_stamp} - {idx}.png" image_path = dirpath / image_name image.save(image_path) print(idx)
2. 显式清理GPU内存(适用于必须动态加载模型的场景)
如果你的业务需要每次调用函数都切换不同模型,那在生成完成后要手动清理无用对象和GPU缓存:
在create函数末尾添加以下代码:
# 删除不再使用的对象 del pipe, scheduler, output # 强制触发Python垃圾回收 import gc gc.collect() # 清空GPU缓存 if torch.cuda.is_available(): torch.cuda.empty_cache() torch.cuda.ipc_collect()
3. 启用diffusers内置内存优化
diffusers提供了多种内存优化工具,能自动管理模型在GPU/CPU之间的调度,减少显存占用的同时,也能让内存更容易被释放:
pipe.enable_model_cpu_offload():适合显存较小的设备,模型层按需加载到GPUpipe.enable_sequential_cpu_offload():轻量化的逐块加载模式pipe.enable_attention_slicing():对注意力层进行切片,降低峰值内存占用
这些方法只需在pipe.to(device)之后调用即可,无需手动管理内存。
内容的提问来源于stack exchange,提问作者M. Özdemir
相关产品推荐
相关产品推荐

