You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Stable Diffusion XL运行报CUDA内存不足,本地与Colab均遇此问题

Stable Diffusion XL CUDA显存不足问题及解决方案

出现的错误如下:

OutOfMemoryError: CUDA out of memory. Tried to allocate 1024.00 MiB. GPU 0 has a total capacty of 14.75 GiB of which 857.06 MiB is free. Process 43684 has 13.91 GiB memory in use. Of the allocated memory 13.18 GiB is allocated by PyTorch, and 602.64 MiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting max_split_size_mb to avoid fragmentation. 

我在使用Stable Diffusion XL生成图像时,已完成Python、PyTorch、Diffusers、CUDA的版本匹配,但仍触发上述显存不足错误。无论在本地搭载NVIDIA RTX 3060(6GB显存)的电脑,还是拥有15GB显存的Google Colab环境中,问题都存在。

已尝试以下方案,但均未解决问题:

  • 未进行模型训练,推理时batch_size设为1;
  • 添加环境变量:PYTHONUNBUFFERED=1;PYTORCH_CUDA_ALLOC_CONF=max_split_size_mb:256;
  • 将生成图像尺寸调整为512x512;
  • 尝试根据RTX 3060与CUDA 11.3版本降级PyTorch至1.8.1,却报错:Could not find a version that satisfies the requirement torch==1.8.1

我的Python代码如下:

from diffusers import DiffusionPipeline, StableDiffusionXLImg2ImgPipeline
import torch
import gc

#for cleaning memory
gc.collect()
del variables
torch.cuda.empty_cache()

model = "stabilityai/stable-diffusion-xl-base-1.0"
pipe = DiffusionPipeline.from_pretrained(
    model,
    torch_dtype=torch.float16,
)
pipe.to("cuda")
pipe.load_lora_weights("model/", weight_name="pytorch_lora_weights.safetensors")

refiner = StableDiffusionXLImg2ImgPipeline.from_pretrained(
    "stabilityai/stable-diffusion-xl-refiner-1.0",
    torch_dtype=torch.float16,
)
refiner.to("cuda")


prompt = "a portrait of maha person 4k, uhd"

for seed in range(1):
    generator = torch.Generator("cuda").manual_seed(seed)
    image = pipe(prompt=prompt, generator=generator, num_inference_steps=25)
    image = image.images[0]
    image.save(f"output_images/{seed}.png")
    image = refiner(prompt=prompt, generator=generator, image=image)
    image = image.images[0]
    image.save(f"images_refined/{seed}.png")

可行解决方案

  1. 模型与精炼器分时加载
    不要同时将base模型和refiner模型加载到GPU中,生成基础图后卸载base模型,再加载refiner进行精炼,用完后及时卸载释放显存。修改代码如下:
from diffusers import DiffusionPipeline, StableDiffusionXLImg2ImgPipeline
import torch
import gc
import os

prompt = "a portrait of maha person 4k, uhd"
output_dir = "output_images"
refined_dir = "images_refined"

# 创建输出目录
os.makedirs(output_dir, exist_ok=True)
os.makedirs(refined_dir, exist_ok=True)

for seed in range(1):
    generator = torch.Generator("cuda").manual_seed(seed)
    
    # 加载base模型生成图像
    pipe = DiffusionPipeline.from_pretrained(
        "stabilityai/stable-diffusion-xl-base-1.0",
        torch_dtype=torch.float16,
    ).to("cuda")
    pipe.load_lora_weights("model/", weight_name="pytorch_lora_weights.safetensors")
    
    image = pipe(prompt=prompt, generator=generator, num_inference_steps=25).images[0]
    image.save(f"{output_dir}/{seed}.png")
    
    # 卸载base模型释放显存
    del pipe
    gc.collect()
    torch.cuda.empty_cache()
    
    # 加载refiner模型精炼图像
    refiner = StableDiffusionXLImg2ImgPipeline.from_pretrained(
        "stabilityai/stable-diffusion-xl-refiner-1.0",
        torch_dtype=torch.float16,
    ).to("cuda")
    
    refined_image = refiner(prompt=prompt, generator=generator, image=image).images[0]
    refined_image.save(f"{refined_dir}/{seed}.png")
    
    # 卸载refiner模型
    del refiner, image, refined_image
    gc.collect()
    torch.cuda.empty_cache()
  1. 启用内存优化与注意力切片
    加载模型时启用内存高效注意力和注意力切片功能,进一步降低显存占用:
pipe = DiffusionPipeline.from_pretrained(
    "stabilityai/stable-diffusion-xl-base-1.0",
    torch_dtype=torch.float16,
    use_memory_efficient_attention=True,
    enable_attention_slicing=True
).to("cuda")
  1. 将部分模型组件移至CPU
    把模型的文本编码器等非核心组件放在CPU运行,仅将UNet放在GPU,适合显存紧张的环境:
pipe = DiffusionPipeline.from_pretrained(
    "stabilityai/stable-diffusion-xl-base-1.0",
    torch_dtype=torch.float16,
    text_encoder=None,
    text_encoder_2=None
).to("cuda")

# 手动加载文本编码器到CPU
from transformers import CLIPTextModel, CLIPTextModelWithProjection
pipe.text_encoder = CLIPTextModel.from_pretrained(
    "stabilityai/stable-diffusion-xl-base-1.0",
    subfolder="text_encoder",
    torch_dtype=torch.float16
).to("cpu")
pipe.text_encoder_2 = CLIPTextModelWithProjection.from_pretrained(
    "stabilityai/stable-diffusion-xl-base-1.0",
    subfolder="text_encoder_2",
    torch_dtype=torch.float16
).to("cpu")
  1. 修正PyTorch版本安装
    RTX3060适配CUDA11.3无需降级到PyTorch1.8,直接安装适配版本即可,执行以下命令:
pip install torch==1.13.1+cu113 torchvision==0.14.1+cu113 --extra-index-url https://download.pytorch.org/whl/cu113

内容的提问来源于stack exchange,提问作者Mahammad Yusifov

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 21:57:50