You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

ComfyUI显存溢出后调用清理接口引发异常求助排查

显存溢出清理后触发CUDA异常的原因分析与解决思路

问题场景

在ComfyUI中执行图生视频、文生视频等显存密集型任务时,频繁遭遇OOM(显存溢出)。为了自动恢复,我在检测到OOM后调用ComfyUI-Easy-Use组件的cleangpu接口清理显存,代码逻辑如下:

if out of memory:
    SystemResourceManager.clear_queue()
    SystemResourceManager.clear_gpu()

@classmethod
def clear_gpu(cls):
    try:
        url = f"{cls.base_url}/api/easyuse/cleangpu"
        req = urllib.request.Request(url, method="POST")
        with urllib.request.urlopen(req, timeout=10) as response:
            if response.status == 200:
                import gc
                gc.collect()
                print("cleargpu success!")
            else:
                print(f"cleargpu failed: {response.status}")
    except Exception as e:
        print(f"cleargpu have error: {e}")

但执行后不仅OOM问题未解决,还触发了新的CUDA异常,完整错误日志如下:

!!! Exception during processing !!! Allocation on device 0 would exceed allowed memory. (out of memory)
Currently allocated     : 21.96 GiB
Requested               : 416.62 MiB
Device limit            : 23.50 GiB
Free (according to CUDA): 31.69 MiB
PyTorch limit (set by user-supplied memory fraction)
                        : 17179869184.00 GiB
Traceback (most recent call last):
  File "/root/ComfyUI/execution.py", line 323, in execute
    output_data, output_ui, has_subgraph = get_output_data(obj, input_data_all, execution_block_cb=execution_block_cb, pre_execute_cb=pre_execute_cb)
  ...(省略中间栈信息)
  File "/usr/local/lib/python3.10/site-packages/diffusers/models/embeddings.py", line 440, in forward
    embeds = embeds + pos_embedding
torch.cuda.OutOfMemoryError: Allocation on device 0 would exceed allowed memory. (out of memory)

Got an OOM, unloading all loaded models.
Exception in thread Thread-2 (prompt_worker):
Traceback (most recent call last):
  ...(重复OOM栈信息)
During handling of the above exception, another exception occurred:

Traceback (most recent call last):
  File "/usr/local/lib/python3.10/threading.py", line 1009, in _bootstrap_inner
    self.run()
  ...(省略中间栈信息)
  File "/usr/local/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1158, in convert
    return t.to(device, dtype if t.is_floating_point() or t.is_complex() else None, non_blocking)
RuntimeError: CUDA error: invalid argument
CUDA kernel errors might be asynchronously reported at some other API call, so the stacktrace below might be incorrect.
For debugging consider passing CUDA_LAUNCH_BLOCKING=1.
Compile with `TORCH_USE_CUDA_DSA` to enable device-side assertions.

关键错误点提取

  1. 初始OOM:显存已用21.96GiB,剩余仅31.69MiB,无法满足416.62MiB的分配请求。
  2. 二次异常:ComfyUI自动执行unload_all_models()卸载模型时,调用model.to(device_to)触发RuntimeError: CUDA error: invalid argument。

核心原因分析

  1. 清理时机与ComfyUI原生逻辑冲突:OOM发生后,ComfyUI会自动启动模型卸载流程(日志中Got an OOM, unloading all loaded models.),此时用户代码同时调用外部cleangpu接口,导致双重内存操作,打乱了ComfyUI对模型状态的追踪,引发后续转移模型时的参数异常。
  2. 外部清理接口与ComfyUI内存管理不兼容:ComfyUI通过model_management模块统一管理模型加载/卸载、显存分配,而cleangpu接口可能直接操作CUDA显存或独立处理模型状态,未与ComfyUI的内部状态同步,导致模型卸载时出现无效参数。
  3. CUDA上下文异常传递:OOM已经导致CUDA上下文处于异常状态,后续的模型转移操作(model.to())会继承该异常状态,触发"invalid argument"错误——这类错误通常是因为之前的CUDA操作失败,导致后续调用的参数或环境无效。
  4. PyTorch内存配置错误:日志中PyTorch limit (set by user-supplied memory fraction)显示为17179869184.00 GiB,这明显是配置错误(远超显卡实际显存),会导致PyTorch的内存限制逻辑失效,加剧OOM风险。

解决建议

  1. 移除自定义OOM清理逻辑:依赖ComfyUI原生的OOM处理机制即可,它会自动执行模型卸载和内存清理,不需要额外调用外部接口。
  2. 同步清理逻辑与ComfyUI执行流程:如果必须自定义清理,需要通过ComfyUI提供的钩子函数(如pre_execute_cb、execution_block_cb)在任务执行间隙清理,避免在OOM异常抛出后并发操作。
  3. 修复PyTorch内存配置:检查启动参数或配置文件,将--gpu-memory-fraction设置为合理值(比如0.9),确保PyTorch的内存限制符合显卡实际容量。
  4. 优化视频生成显存占用:
    • 降低视频分辨率、帧率或时长
    • 启用模型分片(model splitting)或CPU offload
    • 使用fp16精度加载模型

内容的提问来源于stack exchange,提问作者Yiyabo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 05:07:32